Pick an AI answer that recommends vendors in your category and open the sources underneath it. Count how many are yours.
For most brands it’s a short count. McKinsey found that a brand’s own sites make up only 5% to 10% of the sources AI search draws on (McKinsey, October 2025 (opens in a new tab)). In consumer packaged goods and financial services, publishers, user-generated content and affiliate sites made up more than 65% (McKinsey, October 2025 (opens in a new tab)).
You have two audiences now: the person who buys and the AI that decides whether that person ever hears your name. The second one does most of its reading somewhere else. The version of your brand a buyer gets from ChatGPT or Gemini is largely a paraphrase of other people’s pages. Reviews. Forum threads. A roundup someone wrote two years ago. Your competitor’s comparison page.
That makes the answer a job for more than your content team. PR, customer marketing, community and whoever keeps your directory listings current all shape it now. They rarely work from the same list.
A hypothetical source audit
This example is made up, but the pattern will probably look familiar.
You market expense software for mid-size companies. You ask ChatGPT, Gemini, Perplexity and Claude: “What’s the best expense management software for a 300-person company with staff in three countries?” The four answers cite 14 different sources between them.
- Two are yours: the pricing page and an integrations doc nobody has touched in two years.
- Three belong to your biggest competitor: a comparison page that puts the two of you side by side, a blog post on multi-currency reimbursement and a help article.
- Nine belong to third parties: two review sites, a finance publication’s “best tools” roundup, a Reddit thread, a YouTube walkthrough, three software directories and a partner marketplace listing. One directory still shows the per-seat price you dropped last spring. One of the answers repeats it.
- Yours2
- Your competitor3
- Third parties9
None of that shows up in your analytics. The comparison page is the one I’d open first. When the engine explains how you and your rival differ, it’s reading a page your rival wrote.
What the studies say, and why they seem to disagree
The public data is patchy. Most of it comes from SEO and AI visibility vendors, some of whom compete with us. Still, it points one way: the engines lean heavily on a few big third-party platforms, and each engine leans differently.
- In Pew’s March 2025 panel, Wikipedia, YouTube and Reddit were the three most-cited sources in Google’s AI summaries, and together they made up 15% of all the sources listed (Pew Research Center, July 2025 (opens in a new tab)).
- Profound analyzed 680 million citations from August 2024 to June 2025. Among ChatGPT’s ten most-cited sources, Wikipedia accounted for 47.9% of citations. Among Perplexity’s, Reddit accounted for 46.7% (Profound, June 2025 (opens in a new tab)).
- Semrush saw Reddit in close to 60% of ChatGPT responses and Wikipedia in roughly 55% through the summer of 2025. By mid-September both had fallen sharply, to around 10% and under 20% (Semrush, November 2025 (opens in a new tab)). In Google’s AI Mode, LinkedIn showed up in nearly 15% of responses, and the mix stayed far steadier than ChatGPT’s (Semrush, November 2025 (opens in a new tab)).
Read quickly, those numbers clash. Mostly they don’t. Pew’s figure is a share of every source listed. Profound’s are shares of each engine’s ten most-cited sources only, which makes the leaders look bigger. Semrush counts how many responses include a domain at least once. Three denominators, answering three different questions.
Two things hold up across all of them. The sources differ by engine, so a plan built on one engine’s habits leaves gaps on the others. And the mix moves. Semrush caught its drop in the middle of a 13-week study. Ahrefs found that 45.5% of the URLs cited in Google’s AI Overviews were new from one response to the next, even when the answer meant the same thing (Ahrefs, November 2025 (opens in a new tab)).
Third-party pages get two chances at your buyer
The engine reads a review site when it writes the answer. Then the buyer, checking that answer, may land on the same site.
G2’s July 2026 report found review sites (38%) had edged past AI chatbots (37%) as the top source shaping software shortlists (G2, July 2026 (opens in a new tab)). G2 sells reviews, so weigh that. Separately, a 2025 Ahrefs study of 75,000 brands found that brand mentions across the web tracked AI Overview visibility far more closely than backlinks did, with correlations of 0.664 against 0.218 (Ahrefs, May 2025 (opens in a new tab)). That’s correlation, not cause.
So a review profile with six reviews from 2023 hurts you twice, in the answer and in the check that follows it.
Build your source map
Don’t borrow anyone’s map, including a vendor’s. Your category and your buyers’ questions will produce their own. Budget a couple of days the first time, most of it spent reading.
- Pick 10 to 15 buying questions. Category questions that don’t name you, comparisons with named competitors and “alternatives to” questions. Write them the way a buyer types. Build a question set that sounds like your buyers covers this in detail.
- Ask each one on every engine your buyers use, three times, in a fresh session. With that much citation turnover, one pass shows you one moment.
- Log every cited URL. Record the domain, the owner (you, a competitor or a third party), the type (review site, directory, forum, publisher, video, wiki, docs), which engines cited it and how many questions it appeared under. Then note what it says about you: accurate, outdated, wrong, negative or nothing at all.
- Rank by reach, then by damage. Sources cited across several engines and questions go first. Within those, anything wrong or outdated jumps the queue. It’s usually the fastest thing to fix.
- Give every source an action and a named owner. The table below is a starting point.
- Rerun the same questions monthly. Same wording, same engines. Mark which sources are new and which dropped out.
| What’s cited | What to do | Usual owner |
|---|---|---|
| Your page, and the answer favors you | Keep it current and keep the URL stable through the next redesign | Web team |
| Your page, and the answer favors a rival | Say who it’s for, what it costs and why you win, in plain text near the top | Product marketing |
| A competitor’s comparison or blog post | Publish your own fair answer to the same question, with specifics they left out | Product marketing |
| Review sites | Complete the profile, ask real customers for reviews and reply to the critical ones | Customer marketing |
| Directories and marketplaces | Correct pricing, integrations and who the product serves | Web ops or partnerships |
| Forums and communities | Answer as a named employee and say who you work for | Community or DevRel |
| Publisher roundups | Brief the writer, offer original data and ask for factual corrections | PR |
| Wikipedia | Propose fixes on the article’s talk page (its conflict-of-interest rules discourage editing your own article) | Comms |
| Video | Post walkthroughs with accurate titles and transcripts | Content |
Plan for different speeds. You can fix your own page this week. A directory correction might take a support ticket and a few weeks. Reviews and press coverage build over quarters. Report them on separate clocks, or the slow work will look like it’s failing.
And don’t skip your own pages because they’re a small share. They’re the sources you fully control, and they’re often harder for machines to read than they look. Adobe scored the average US retail product page at 66% on its AI content visibility measure, where 50% means half the content isn’t machine-readable (Adobe, April 2026 (opens in a new tab)). That’s retail, measured with Adobe’s own tool. The lesson travels anyway. A price baked into an image, or a spec sheet that only loads after a click, is easy for a person to see and easy for a crawler to miss.
Where this goes wrong
Counting a citation as a win. A cited page can be arguing against you. Track the sources that say something wrong or negative separately, or your map will flatter you.
Trusting the source list too much. In a 2025 study of news answers led by the BBC and the European Broadcasting Union, a third of responses had serious sourcing problems (EBU, October 2025 (opens in a new tab)). That was news, not brand questions. It’s still a reason to read the cited page before you act on it.
Faking it. Planted reviews and sock-puppet forum posts break the platforms’ rules. They’re also exactly what a skeptical buyer finds when they go checking.
Chasing one engine’s favorite domain. The Semrush numbers show how quickly a favorite can lose its place. Work on the sources that turn up across engines first.
Keeping the map current
Do the first map by hand. Reading twenty review pages and forum threads about your own product is humbling, and it shows you exactly where the engines get their lines about you.
Keeping it current is the tedious part: four engines, a dozen questions, repeat runs, every month. Contentstack Canoe does that part for you. It takes the sources each engine actually cites and sorts them by owner (your domain, competitor domains or third parties). Each run flags the sources that are new and the ones you lost. Its citation recommendations stick to channels you can earn, like your own site, review sites, directories, communities, PR and partners. They never point you at a competitor’s property.
Start with the three domains that show up under the most answers. There’s a fair chance nobody on your team has looked at one of them this year.
