What are AI crawlers? Definition, user agents and how to control them

AI crawlers are the bots AI companies use to read web pages for model training, AI search indexes and live answers. Main user agents and how to control them.

  • Alvaro Peña de Luna Alvaro Peña de Luna
  • date icon

    Wednesday, Sep 30, 2026

AI crawlers are automated bots that AI companies send to fetch web pages, either to collect training data for their models, to build the search index their assistants cite, or to read a specific page when a user asks about it. Each one identifies itself with a user agent that you can allow or block in robots.txt.

Knowing which is which matters, because the same robots.txt line can keep you out of a future model or out of the answers an assistant gives today.

How AI crawlers work

Technically, AI crawlers behave like search engine crawlers: they request a URL, read the HTML and follow links. What changes is what happens to the page afterwards. AI companies run crawlers with three different jobs:

  • Training crawlers collect public pages that may be used to train future models. Their effect is slow: your content only shapes what a model knows once a new model is trained on it.
  • Search crawlers build the index an assistant searches when it answers with web results. This is the crawl that decides whether your pages can be retrieved and cited.
  • User-triggered agents fetch a page on demand, during a conversation, when a user asks the assistant to read or check something.

Some companies also publish robots.txt tokens that have no crawler of their own. They ride on an existing crawl and only tell the company whether it may use your content for its AI products.

The main AI crawler user agents

User agent token Company Job
GPTBot OpenAI Training
OAI-SearchBot OpenAI ChatGPT search index
ChatGPT-User OpenAI Visits when a user asks
ClaudeBot Anthropic Training
Claude-SearchBot Anthropic Search index
Claude-User Anthropic Visits when a user asks
PerplexityBot Perplexity Search index
Perplexity-User Perplexity Visits when a user asks
Google-Extended Google Token only: controls use in Gemini models
CCBot Common Crawl Open dataset used for training

Each vendor documents its own list, and it changes: see OpenAI's crawlers and Google's common crawlers. Note that Google AI Overviews and AI Mode use Googlebot's regular crawl, not Google-Extended, so blocking Google-Extended does not take you out of them.

Why AI crawlers matter for brands

An AI engine can only cite what it can read. If its search crawler is blocked, by robots.txt or by a firewall rule, your pages cannot be retrieved for its answers, and the engine will describe your brand from other people's pages instead: review sites, forums, competitors' comparisons. Your product pages, pricing and documentation stop being the source.

Training crawlers work on a longer horizon. What a model knows about your brand without searching comes from what it was trained on, up to its knowledge cutoff. Blocking training is a legitimate choice, especially for publishers, but it is a trade-off, not a free privacy setting.

How to control AI crawlers in robots.txt

A common policy is to stay visible in AI search while opting out of model training: allow the search and user-triggered crawlers and disallow the training ones.

User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /

Two rules decide how crawlers read this. A crawler that has its own group follows only that group and ignores the User-agent: * rules, so repeat there anything it must also respect. Within a group, the longest matching path wins. Before you publish a change, paste the edited file into the free robots.txt tester and test the URLs that matter against each bot.

How to check which AI crawlers reach your site

  1. Check access. The AI crawler checker reads your robots.txt the way each crawler does, requests the page with the main crawlers' user agents to catch firewall and CDN blocks, and flags noindex or nosnippet directives.
  2. Check real visits. The log file analyzer reads your access log in your browser and shows which AI crawlers came, which pages they read and where they got errors.
  3. Check the result. Access is the floor, not the goal. Mencoro tracks how ChatGPT, Perplexity, Google AI Overview and Google AI Mode mention your brand next to your competitors, so you can see whether opening the door changes the answers.

Common mistakes

  • One switch for every AI bot. CDN options that block "AI bots" in one click usually stop search and user-triggered crawlers too, which removes you from the answers.
  • Blocking GPTBot to leave ChatGPT search. ChatGPT search relies on OAI-SearchBot. GPTBot is about training.
  • Blocking Google-Extended to leave AI Overviews. It does not work that way. To limit how Google shows a page in its AI features, use snippet directives such as nosnippet.
  • Trusting the user agent. Anyone can claim to be GPTBot. For a strict audit, check requests against the IP ranges each company publishes.
  • Hiding content behind JavaScript. Do not assume every AI crawler renders pages the way Googlebot does. Serving the main content in the HTML is the safe choice.

For every AI crawler with its user agent, IP ranges and robots.txt behaviour, see the AI crawler index.

FAQ

Frequently asked questions

A training crawler, such as GPTBot or ClaudeBot, collects pages that may be used to train future models. A search crawler, such as OAI-SearchBot or PerplexityBot, builds the index an assistant searches when it answers with web results. Blocking the first keeps your content out of training; blocking the second keeps your pages out of the answers.
Not entirely. Blocking search and user-triggered crawlers stops those engines from reading and citing your pages, but they can still name your brand from what the model learned in training or from other sites that write about you. You lose control of the source, not the mention.
The main AI companies document how their crawlers read robots.txt, but robots.txt is a voluntary standard and anyone can fake a user agent. If you need to enforce a block, do it in your firewall or CDN and verify requests against the IP ranges each vendor publishes.
Look for their user agents in your server access logs. The free Mencoro log file analyzer reads the log in your browser and groups every bot by what it is for, with its requests, the pages it reads and the errors it gets.

Start tracking your brand in AI search today

Monitor how AI engines cite your brand, track keyword positions, and benchmark against competitors, all in one platform.