Take Tarnwell Cycles, a made-up bike brand we use for examples. In March its team launches a page for a new commuter e-bike, the Tarnwell Metro. It’s a good page. It ranks on Google within a few weeks, and its range and weight figures beat what Aldren and Solvik publish.

Six weeks later, ChatGPT and Perplexity still recommend Aldren and Solvik when someone asks about commuter e-bikes. When they cite Tarnwell at all, it’s a blog post from two years ago.

Two things are going on, and nobody on the marketing team knows about either. Last year, after a scraping incident, the security team turned on a CDN rule that challenges bots it doesn’t recognize. And the spec table on the new page is loaded by JavaScript after the page opens. So some AI crawlers get a challenge page, and the ones that get through see a product page with no specs on it.

Nobody did anything wrong. Security stopped the scraping and the web team built a fast page. The gap is that nobody looked at the page the way an AI crawler sees it. This article is about that step: whether the engines can reach and read your important pages at all, before you worry about what they say about you.

Every engine sends more than one bot

People say “AI crawlers” as if it’s one thing. Most engines run several, and each does a different job:

  • Search crawlers build the index the engine searches when it answers with web results.
  • User fetchers visit a page because someone asked the assistant about it, right now.
  • Training crawlers collect pages that may be used to train future models.

Here’s how the three vendors with public crawler docs name theirs:

AI crawlers by job, as each vendor documents them
EngineSearchUser fetchTraining
ChatGPTOAI-SearchBotChatGPT-UserGPTBot
ClaudeClaude-SearchBotClaude-UserClaudeBot
PerplexityPerplexityBotPerplexity-UserNone named
Names from each vendor’s crawler docs, checked in September 2026. Perplexity says PerplexityBot isn’t used to train AI models.

Blocking each one does something different. OpenAI says sites that opt out of OAI-SearchBot won’t be shown in ChatGPT search answers, though they can still appear as navigational links, and that a robots.txt change takes about a day to apply (OpenAI (opens in a new tab)). Anthropic says blocking Claude-SearchBot may reduce your visibility and accuracy in Claude’s search results (Anthropic (opens in a new tab)). Perplexity describes PerplexityBot as the crawler that surfaces and links sites in its results (Perplexity (opens in a new tab)).

Blocking the training crawlers is a policy decision, and a reasonable one to make with your legal team. Blocking the search crawlers is a different decision, because it takes you out of the answers. It’s easy to treat them as one thing. Plenty of sites block AI crawlers across the board: Cloudflare found AI crawlers were the user agents most often fully disallowed in robots.txt files in 2025 (Cloudflare, December 2025 (opens in a new tab)).

Google works a little differently. AI Overviews and AI Mode are built on Google Search, and Google says a page needs to be indexed and eligible for a snippet to show up as a link in them (Google Search Central (opens in a new tab)). So for Google, the question is the one your SEO team already asks.

Blocked without knowing it

A common surprise is a robots.txt that says yes while something else says no.

That something is usually a CDN or firewall rule. Security teams add bot rules for good reasons: scrapers, credential stuffing, traffic spikes during a launch. Many of those rules challenge or block traffic that doesn’t look like a person in a browser, and AI crawlers don’t. The rule doesn’t show up in robots.txt, and nobody in marketing gets told.

It’s also becoming a default. Since July 2025, Cloudflare, which says it handles traffic for 20% of the web, has asked every new domain at sign-up whether to allow AI crawlers, and it blocks them by default (Cloudflare, July 2025 (opens in a new tab)). That’s a fine choice for a lot of sites. It’s worth knowing whether yours made it on purpose.

The logs will tell you. Look in your server or CDN logs for the AI user agents. A crawler that keeps getting a 403 (forbidden), a 429 (too many requests) or a 200 with a challenge page instead of your content is being blocked, whatever robots.txt says.

When you allow them back in, don’t trust the name alone. Anyone can put “PerplexityBot” in a request. Perplexity suggests allowlisting its crawlers by user agent and by its published IP ranges together (Perplexity (opens in a new tab)), and that’s a good habit for every engine that publishes them.

Reachable but empty

The second problem is harder to spot. The crawler gets in, gets a 200 and still misses the part of the page that matters.

Vercel looked at crawler traffic across its network and found that none of the major AI crawlers it saw rendered JavaScript. That included OpenAI’s three crawlers, ClaudeBot and PerplexityBot. Some of them download JavaScript files, but they don’t run them. Gemini was the exception, because it uses Googlebot’s infrastructure, which does render pages (Vercel, December 2024 (opens in a new tab)).

So if your prices, specs, comparison tables or FAQ answers are added to the page by JavaScript after it loads, most AI crawlers get the page without them. Text that’s in the HTML and only hidden by CSS, like a closed tab or accordion, is usually fine. Reviews widgets, pricing calculators and product configurators often aren’t.

The same hypothetical page before and after JavaScript runs. Tarnwell Cycles is a fictional brand.

That Vercel data is from late 2024 and from one network. Crawlers change, and some may start rendering pages. I wouldn’t plan around it, though, because the check is cheap: turn JavaScript off in your browser and reload the page. If the spec table disappears, it isn’t there for most AI crawlers either.

Readable but restricted

Then there are pages that are reachable and readable but carry an instruction that limits what an engine can do with them. These usually get there by accident: a template flag left over from staging, a plugin setting or a header someone added at the CDN.

  • noindex in a robots meta tag or an X-Robots-Tag header keeps the page out of the index altogether.
  • nosnippet, max-snippet and data-nosnippet limit how much of the page can be shown. Google says these apply to its AI features too (Google Search Central (opens in a new tab)).
  • A canonical that points somewhere else tells engines another URL is the real one. If every product variant points at the category page, the variant pages drop out.
  • A missing or stale sitemap won’t block anything, but new pages take longer to be found. Declare it in robots.txt, list your key pages and keep the dates current.

Headers are the easy ones to miss, because you won’t see them when you view the page source. Check them in your browser’s network tab or with a command-line fetch.

What matters less than you’ve heard

This is where the advice gets noisy, and where the sources don’t agree.

A lot of articles tell you to add an llms.txt file. It’s a proposal from September 2024 for a markdown file that points AI tools at your most useful pages (llmstxt.org (opens in a new tab)). It’s cheap to add and I don’t think it hurts. But Google says you don’t need new machine-readable files, AI text files or special structured data to appear in its AI features (Google Search Central (opens in a new tab)), and I haven’t seen the other major engines say they rely on it.

My read: if you only have an afternoon, spend it on the access checks below. An llms.txt file won’t help a page your firewall blocks, or a page that shows an empty table without JavaScript. Once those are fixed, add one if it’s easy.

Structured data is similar. It helps engines understand a page and it’s worth having for the usual search reasons. It won’t make up for content the crawler never receives.

The AI crawler access check

You can do this in an afternoon with a browser, a terminal and someone from security. Start with five pages: the homepage, your main product page, pricing and your two most important comparison or category pages.

  1. Read robots.txt as each bot. Check the rules for OAI-SearchBot, Claude-SearchBot, Claude-User and PerplexityBot. OpenAI (opens in a new tab) and Perplexity (opens in a new tab) say their user fetchers (ChatGPT-User and Perplexity-User) may not follow robots.txt, because a person asked for the page, so for those your firewall rules decide. Decide on the training crawlers separately, with legal.
  2. Ask security what’s switched on. Bot management, challenge pages, rate limits and country blocks at the CDN or firewall. Get the list in writing.
  3. Fetch each page as each search bot. Send the bot’s user agent and compare the response with a normal browser fetch. Firewalls that check bot IP addresses can treat the real crawler differently from your check, so a pass is a good sign, not proof.
  4. Reload key pages with JavaScript off. Pricing, product and comparison pages first. Anything that disappears needs to be in the HTML.
  5. Check directives and headers. Look for noindex, nosnippet and max-snippet in meta tags and X-Robots-Tag headers, and make sure each page’s canonical points at itself.
  6. Check the sitemap. It should be declared in robots.txt, list your key pages and have current dates.
  7. Read 30 days of logs. Which AI crawlers hit which pages, and which status codes they got. Clusters of 403s and 429s show you where to look.
  8. Add it to the release checklist. Run it again after CDN changes, redesigns and migrations.

For step 3, a terminal is enough to catch the obvious cases. This compares what a browser and a search bot get, then checks whether the spec table is in the HTML before any JavaScript runs:

Terminal
# status code for a browser, then for a search bot
curl -s -o /dev/null -w "%{http_code}\n" https://tarnwell.example/bikes/metro
curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot" https://tarnwell.example/bikes/metro

# is the spec table in the raw HTML?
curl -s https://tarnwell.example/bikes/metro | grep -c "Range"

Write down what you find for each page and each bot. The next check goes much faster when you can compare it with the last one.

Keep it checked after every release

Access problems don’t arrive on their own. They come with other work: a new bot rule after an incident, a redesign that moves specs into a JavaScript component, a migration that ships the staging robots.txt. None of those changes is about AI search, so nobody thinks to check.

That’s why I’d put this on the release checklist rather than the content calendar. After the first time, it takes a few minutes.

We built a site audit into Contentstack Canoe for the same reason. On every plan, it runs automatically with your monitor runs. It checks your robots.txt rules for AI search, answer-time and training crawlers. It also checks whether CDN or firewall rules block or challenge AI crawlers that a browser gets past and whether your homepage has noindex, nosnippet or canonical problems. It also crawls a sample of your pages and flags broken pages, missing or duplicate titles, heading problems, thin content and slow server responses. Each finding says why it matters and how to fix it.

It doesn’t replace your logs. It can only flag likely blocks, because a firewall that checks the engines’ own IP addresses may treat the real crawlers differently from our check.

Either way, start with the page you’d most want an engine to quote, and load it once with JavaScript off.