Here’s a hypothetical. An expense software company wants to see how it shows up in AI answers. Someone exports the top 20 terms from the keyword tool and runs them through ChatGPT: “expense management software,” “best expense app,” “expense report tool.” The brand is named in about half the answers. Everyone’s pleased.
Then a sales rep reads the list and says none of her buyers this quarter asked anything like it. They asked: “We’re 120 people, most of them in the field, and finance is two people. What should we use so receipts stop going missing, and does it sync with NetSuite?”
Ask an assistant that and you may get a different set of brands, without yours in it.
The question set is the instrument. Visibility, share of voice and average position only describe the questions you chose. Choose badly and the report is still precise and repeatable. It’s just wrong the same way every month, which is harder to spot than noise.
Keywords and questions measure different demand
A keyword is what’s left of a question after someone squeezes it into a search box. Keyword volume tells you how often people typed a string into Google, and very little about what they ask an assistant.
The full question keeps the details that decide the answer: team size, region, the tool they’re leaving, the system it has to connect to, the budget someone upstairs set. A tool that’s right for 20 people can be wrong for 2,000, and a good answer will say so.
Head terms also flatter whoever is already biggest. “Best expense software” mostly returns the brands mentioned everywhere. Ask the specific question and a smaller brand that fits has a real shot.
And buyers bring decision questions to AI. In G2’s July 2026 research, buyers used AI chatbots most to understand total cost of ownership (51%) and build shortlists (51%) (opens in a new tab). A keyword export is heavy on “what is” and light on “what will this cost us in year two.”
There’s no Search Console for ChatGPT. The prompt logs stay with the companies that run the engines, so you rebuild buyer questions from other evidence.
Where real buyer questions live
The raw material is already inside your company, just not in the SEO team’s tools.
- Discovery calls. The first ten minutes, where the buyer explains what they’re trying to fix, are the closest thing you have to what they typed into an assistant the week before.
- Lost-deal notes. Which comparisons you lose, and on what.
- Support and onboarding tickets. New customers ask about whatever worried them before they signed.
- Review sites. The “what do you dislike” answers, yours and your competitors’, are the objections buyers take to an engine.
- Community threads. A Slack or Reddit post asking for a recommendation is already shaped like a prompt: context, constraints, then the ask.
- RFPs and security questionnaires. Decision-stage questions in the buyer’s own words.
A keyword is what’s left of a question after someone squeezes it into a search box.
Keep the keyword tool for vocabulary: the words buyers use for your category and the names your competitors get searched alongside. Use it as a dictionary, not a script.
One persona is one seat at the table
It’s tempting to write a question set for one imagined buyer. Forrester’s 2026 State of Business Buying research found that, on average, 13 internal stakeholders and nine external participants influence a B2B purchase (opens in a new tab).
Each asks something different, and each question can produce a different shortlist. For the hypothetical expense company:
- The controller wants approval rules by department and an audit trail the auditors will accept.
- A field sales manager wants the simplest app for reps who photograph receipts on their phones.
- IT wants single sign-on and users added and removed automatically.
- The CFO wants to know what this usually costs per user at 120 people.
You can be the obvious answer for the field manager and missing for the controller. A set written for one persona will never show you that.
Balance the set on purpose
Change the mix and the number changes, so choose the mix deliberately.
Most questions shouldn’t name you. Category questions (“what should we use for expenses?”) are where buyers meet brands they’ve never heard of. In G2’s April 2026 survey on chatbots and software buying, 33% of buyers said they’d bought from a vendor they didn’t know before (opens in a new tab).
Watch for the vanity set. If half your questions contain your brand name, visibility will look great, because an engine asked about you will talk about you. Branded questions still matter. They show how you’re described and whether the facts are right. They can’t tell you if anyone finds you.
Include head-to-heads and switchers. A comparison with a named competitor shows how the engine frames the choice. “Alternatives to [competitor]” questions catch buyers who are already unhappy with someone else.
A reasonable start: make category questions that never name you the biggest group, then add a smaller share about you, a few head-to-heads and a couple of “alternatives to” questions. Adjust it once you have data.
- MostCategoryNever name you
- SomeBrandedAbout you
- A fewComparisonWith a named competitor
- A coupleAlternatives“Alternatives to [competitor]”
Tag every question by stage: awareness, consideration or decision. OpenAI says ChatGPT is where discovery, consideration and decision-making now happen together (opens in a new tab). Cover all three and read visibility stage by stage, because missing at awareness and missing at decision need different fixes.
Write each question the way a person types it
Five rules I’d hold every question to:
- First person, full sentence. “We’re a 40-person agency looking for…” beats “agency project management tool.”
- One to three real constraints. Size, region, stack, budget or deadline. Nobody types all of them.
- Their vocabulary, not yours. If you say spend management platform and buyers say expense app, write expense app.
- No leading questions. “Why is [your brand] the best choice for…” measures your prompt writing, nothing else.
- The right language and market. Ask in your buyers’ language, about the market they buy in. An expense question from Munich and one from Texas should get different answers.
The same hypothetical company, rewritten:
| Keyword tool gives you | A buyer actually asks |
|---|---|
| expense management software | We’re 120 people, mostly in the field, with a two-person finance team. What should we use so receipts stop going missing? |
| [Competitor] pricing | Is [Competitor] worth it at our size, or is there something cheaper that still handles mileage? |
Where the data disagrees
How far to lean into long, assistant-style questions depends on your buyers. The research doesn’t settle it.
G2 found in April 2026 that 51% of B2B software buyers start research in an AI chatbot more often than in Google (opens in a new tab). G2 sells reviews, so it has a stake in that story. A Gartner survey of 377 US consumers, run in mid-2025, found only about a third think generative AI chatbots are as effective as search engines (opens in a new tab). Different people, different questions, most of a year apart. Neither is wrong. They describe different buyers.
My read: if you sell software to businesses, weight the set heavily toward full, assistant-style questions. If you sell to consumers, keep a slice of shorter queries for Google’s AI Overviews, which Google says reach more than 2.5 billion people a month (opens in a new tab). Set the split from your own buyers, not someone else’s survey.
Keep it still long enough to learn something
The engines change on their own schedule. Semrush’s tracking showed ChatGPT citing Reddit in close to 60% of responses in early August 2025 and around 10% by mid-September (opens in a new tab). Semrush sells AI visibility tools, so treat that as vendor data, but a break that sharp is hard to miss. Rewrite your questions that month and you can’t tell if your numbers moved because of you or the engine.
So freeze the set. Give it a version number and a date, and keep it fixed for at least a quarter. Add new questions as a separate batch rather than editing old ones, and never draw one trend line across two versions.
Ask each question more than once, too. SparkToro found less than a 1 in 100 chance (opens in a new tab) that ChatGPT or Google’s AI returns the same list of brands twice, and suggests measuring visibility across dozens to hundreds of prompts, each run several times. Judging a score built that way is its own piece: Can you trust an AI visibility score?
A checklist for Monday
- Collect 30 real questions from the last 90 days of call notes, lost deals, tickets, reviews and community threads. Keep the buyer’s words.
- Pick one product per set. Mixed products give you a number nobody can act on.
- Group by role and tag by stage. At least three roles from the buying group. Every question marked awareness, consideration or decision.
- Balance the mix: mostly category questions, plus branded, comparison and “alternatives to” questions.
- Rewrite into prompt shape. First person, one to three constraints, buyer vocabulary, right language and market.
- Strip the bias. No brand name in category questions and no positioning adjectives anywhere.
- Choose competitors from real deals, not from the list you wish you competed against.
- Show the list to two sales reps. Cut anything neither has heard this month.
- Freeze it and version it. Review it once a quarter.
Let a tool draft it, then argue with the draft
This is slow the first time. That’s fine. Arguments about which questions belong are usually arguments about who you sell to, and you want those before the numbers arrive.
Contentstack Canoe drafts a first set along these lines. It reads your site and writes questions from your products, personas and named competitors, using live research on what buyers in your space ask. The set is balanced across category, branded, comparison and alternatives questions. Each question is tagged by stage and you pick the region and language for each monitor. On the Growth plan you can edit, add or remove any question, so run the checklist on the draft before the first scheduled run. The report also shows which personas and products no question covers yet, which is a good place to start your next batch.
Before any of that, talk to sales. Ask two reps what their last three buyers wanted to know before price came up, and write it down word for word.
