Here’s a hypothetical Monday. Your AI visibility score went from 61 to 57 over the weekend. Someone in the leadership channel has noticed and wants to know what happened.
You could build a story: a competitor launched something, a review site reshuffled its rankings, the new pricing page hurt you. Or you could ask how many answers it took to move the number four points.
Start with the second. If that score is a visibility rate built from 20 questions across four engines, four points is about three answers out of 80. On engines that rarely say the same thing twice, three answers is weather.
Every AI visibility number is a sample, including the ones we produce. ChatGPT, Gemini, Perplexity and Claude write a fresh answer every time you ask. You’re estimating how often they name you from a limited number of tries, and how far you can trust that estimate depends on how it was made.
Most teams are only now picking their first number. In late 2025, McKinsey found that just 16% of brands systematically tracked AI search performance (McKinsey, October 2025 (opens in a new tab)). Whatever shows up first tends to become the baseline everyone argues from.
Why the same question gets a different answer
In January 2026, SparkToro had 600 volunteers run the same 12 prompts through ChatGPT, Claude and Google’s AI, 2,961 runs in all. For ChatGPT and Google’s AI, there was less than a 1 in 100 chance that two responses named the same list of brands, and about a 1 in 1,000 chance of getting the same list in the same order (SparkToro, January 2026 (opens in a new tab)).
Ahrefs found something that sounds like the opposite. Across more than 43,000 keywords, AI Overview content changed every 2.15 days on average, and about 45% of cited URLs were new from one response to the next. Yet the meaning barely moved: consecutive answers had an average semantic similarity of 0.95 out of 1 (Ahrefs, November 2025 (opens in a new tab)).
Both can be true. The exact output is unstable, while the tendency underneath it holds fairly steady. SparkToro’s Rand Fishkin draws the same line: visibility percentage, measured across many prompts run many times, is a reasonable metric, and ranking position in AI answers isn’t. He also says his study wasn’t peer-reviewed and calls for bigger ones. Fair enough. It’s still the most useful public test of the question I’ve seen.
There’s a second source of movement, and it has nothing to do with you. Semrush tracked citations in ChatGPT, AI Mode and Perplexity over 13 weeks in 2025. In early August, Reddit showed up in about 60% of ChatGPT responses and Wikipedia in about 55%. By mid-September, Reddit had fallen to around 10% and Wikipedia to about 20%, while AI Mode and Perplexity stayed relatively stable (Semrush, November 2025 (opens in a new tab)). Semrush sells a competing product and those figures are read off its charts. The point survives anyway: sometimes an engine changes how it picks sources, and your number moves with it.
So a score can move for three reasons: sampling noise, an engine update or a real change in what the engines read about you. Only the third is yours to act on, and a single number won’t tell you which one you’re looking at.
Four layers behind every score
When I look at a scoring method, ours or anyone else’s, I check four layers in this order.
-
What was asked. Which questions, on which engines, how many times. This is the sample, and it caps the quality of everything built on it. Watch for question sets that name your brand. If half the questions read like “Is [your brand] good for mid-size teams?”, you’ll be visible in at least half the answers by construction.
The questions that teach you something are the ones where the buyer doesn’t know you yet, like “What should a 40-person agency use for project billing?” (More on building that set in Build a question set that sounds like your buyers, not your keyword tool.)
-
How it was counted. A mention is either matched in code or judged by a model reading the answer. Code is boring and repeatable: the same answer gets the same count every time. A model deciding “was the brand mentioned?” adds its own variance on top of the engine’s. Ask the same about citations. Were they read from the sources the engine attached to its answer, or inferred from the text afterward?
Code has blind spots too. Take a hypothetical brand called Harbor. A simple text match will credit it every time an answer mentions a safe harbor clause. Nobody catches that without reading the answers.
-
What was judged. Sentiment, “recommended versus merely mentioned” and perceived market position are all one model’s read of another model’s answer. They’re useful, and they’re opinions. Each should be labeled as an opinion and shown next to the quote it came from, so you can disagree.
-
How it was combined. Most tools roll everything into one composite. Ask what’s in it, how it’s weighted and what happens when the engines disagree. If ChatGPT names you in 80% of answers and Perplexity in 20%, a blended 50 describes neither. And if the vendor can’t tell you what’s in the composite, you won’t be able to explain why it moved.
One question cuts across all four: does the method stay the same from run to run? When Ahrefs updated its study of AI Overview citations in March 2026, the share of cited pages that also appeared in Google’s first ten results fell from about 76% to 38%, and Ahrefs said part of that drop came from improving how it detects citations (Ahrefs, March 2026 (opens in a new tab)). BrightEdge, measuring much the same thing, put the overlap at about 17% (BrightEdge, February 2026 (opens in a new tab)). Two careful teams, one question, very different numbers. The method is part of the number.
(Ahrefs and BrightEdge sell tools in this space too. So do we. Read every vendor study with that in mind.)
The noise test: seven checks for Monday
This takes a spreadsheet, the raw answers and about two hours.
Count the answers behind the number. Multiply questions by engines by runs. With 20 questions on four engines, one answer flipping moves that engine’s visibility by 5 points and the four-engine average by 1.25. Write that down. It’s the smallest step your number can take.
Find your noise floor. Compare two consecutive runs where nothing changed on your side: no launches, no new pages, no PR push. The gap between them is your noise floor. Treat any move inside it as noise until the next run confirms it.
Audit ten answers by hand. Pick ten at random. Check what the tool recorded (named or not, position and cited URLs) against the full answer text. If it gets two of the ten wrong, your score carries that error.
Read those same answers for accuracy. A mention count can’t tell you whether what the engine said about your pricing is true. When journalists at 22 public broadcasters reviewed more than 3,000 answers from ChatGPT, Copilot, Gemini and Perplexity, almost half had at least one significant issue (EBU, October 2025 (opens in a new tab)). That study covered news questions (not brand questions), and I wouldn’t bet on your category faring much better.
Strip out the questions that name you. Recompute visibility on what’s left. If it drops sharply, you know what was holding it up.
Split it by engine. Put the per-engine numbers next to the blended one. If one engine is carrying the average, that engine is the story.
List the judgment calls. Mark every figure that comes from a model’s read. Keep those out of any sentence that starts with “we’re up” or “we’re down” unless the counted numbers agree.
Reporting it upstairs without overclaiming
Your buyers already treat AI answers with some suspicion. In TrustRadius’s 2026 survey of 1,862 technology buyers, 94% of those who used AI said they fact-check its responses at least some of the time (TrustRadius, July 2026 (opens in a new tab)). Give your CMO the same chance with your score.
A few habits help:
- Report the trend across three or more runs. One run is an anecdote with a decimal point.
- Drop the decimal. A 57.3 implies precision that 80 answers can’t support.
- Say what it’s built on. “57, from 80 answers across four engines” lets people judge the move for themselves.
- Show one real answer. Put a verbatim answer next to the chart, ideally one where a competitor wins.
- Keep counted and judged apart. Visibility and citations first. Sentiment second, labeled as a model’s read.
Holding our own number to the same test
It’s only fair to run our own product through this. In Contentstack Canoe, each question goes to every engine on your plan (ChatGPT, Gemini, Perplexity and Claude on paid plans) exactly as a buyer would type it, with web search on. Mentions are counted in code, not judged by a model. Citations are the links each engine actually showed, not a guess. Unless you change the question set, every run uses the same questions, engines and scoring, so this month compares with last month. The overall score combines the engines so one outlier can’t swing it. Sentiment and the market leaderboard are a model’s read, and we say so.
None of that makes the engines deterministic. Nothing can. A single report is still one sample, so the trend is what you should read. What the method does is hold our side of the measurement fixed, so when the number moves, the answers moved, not the ruler. You can open any question and read each engine’s full answer, which is where the ten-answer audit starts.
Run the noise test on whatever you use now, ours included. Then decide how much weight the number can carry.
