AI crawlers are automated bots that AI companies send to fetch web pages, either to collect training data for their models, to build the search index their assistants cite, or to read a specific page when a user asks about it. Each one identifies itself with a user agent that you can allow or block in robots.txt.
Knowing which is which matters, because the same robots.txt line can keep you out of a future model or out of the answers an assistant gives today.
How AI crawlers work
Technically, AI crawlers behave like search engine crawlers: they request a URL, read the HTML and follow links. What changes is what happens to the page afterwards. AI companies run crawlers with three different jobs:
- Training crawlers collect public pages that may be used to train future models. Their effect is slow: your content only shapes what a model knows once a new model is trained on it.
- Search crawlers build the index an assistant searches when it answers with web results. This is the crawl that decides whether your pages can be retrieved and cited.
- User-triggered agents fetch a page on demand, during a conversation, when a user asks the assistant to read or check something.
Some companies also publish robots.txt tokens that have no crawler of their own. They ride on an existing crawl and only tell the company whether it may use your content for its AI products.
The main AI crawler user agents
| User agent token | Company | Job |
|---|---|---|
GPTBot | OpenAI | Training |
OAI-SearchBot | OpenAI | ChatGPT search index |
ChatGPT-User | OpenAI | Visits when a user asks |
ClaudeBot | Anthropic | Training |
Claude-SearchBot | Anthropic | Search index |
Claude-User | Anthropic | Visits when a user asks |
PerplexityBot | Perplexity | Search index |
Perplexity-User | Perplexity | Visits when a user asks |
Google-Extended | Token only: controls use in Gemini models | |
CCBot | Common Crawl | Open dataset used for training |
Each vendor documents its own list, and it changes: see OpenAI's crawlers and Google's common crawlers. Note that Google AI Overviews and AI Mode use Googlebot's regular crawl, not Google-Extended, so blocking Google-Extended does not take you out of them.
Why AI crawlers matter for brands
An AI engine can only cite what it can read. If its search crawler is blocked, by robots.txt or by a firewall rule, your pages cannot be retrieved for its answers, and the engine will describe your brand from other people's pages instead: review sites, forums, competitors' comparisons. Your product pages, pricing and documentation stop being the source.
Training crawlers work on a longer horizon. What a model knows about your brand without searching comes from what it was trained on, up to its knowledge cutoff. Blocking training is a legitimate choice, especially for publishers, but it is a trade-off, not a free privacy setting.
How to control AI crawlers in robots.txt
A common policy is to stay visible in AI search while opting out of model training: allow the search and user-triggered crawlers and disallow the training ones.
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /
Two rules decide how crawlers read this. A crawler that has its own group follows only that
group and ignores the User-agent: * rules, so repeat there anything it must also respect.
Within a group, the longest matching path wins. Before you publish a change, paste the edited file
into the free robots.txt tester and test the URLs that matter
against each bot.
How to check which AI crawlers reach your site
- Check access. The AI crawler checker reads your robots.txt the way each crawler does, requests the page with the main crawlers' user agents to catch firewall and CDN blocks, and flags noindex or nosnippet directives.
- Check real visits. The log file analyzer reads your access log in your browser and shows which AI crawlers came, which pages they read and where they got errors.
- Check the result. Access is the floor, not the goal. Mencoro tracks how ChatGPT, Perplexity, Google AI Overview and Google AI Mode mention your brand next to your competitors, so you can see whether opening the door changes the answers.
Common mistakes
- One switch for every AI bot. CDN options that block "AI bots" in one click usually stop search and user-triggered crawlers too, which removes you from the answers.
- Blocking GPTBot to leave ChatGPT search. ChatGPT search relies on OAI-SearchBot. GPTBot is about training.
- Blocking Google-Extended to leave AI Overviews. It does not work that way. To limit
how Google shows a page in its AI features, use snippet directives such as
nosnippet. - Trusting the user agent. Anyone can claim to be GPTBot. For a strict audit, check requests against the IP ranges each company publishes.
- Hiding content behind JavaScript. Do not assume every AI crawler renders pages the way Googlebot does. Serving the main content in the HTML is the safe choice.
For every AI crawler with its user agent, IP ranges and robots.txt behaviour, see the AI crawler index.
Related glossary terms
- GPTBot: OpenAI's training crawler, in detail.
- llms.txt: a curated map of your key pages for AI models.
- Retrieval-augmented generation (RAG): how engines use the pages search crawlers collect to write answers.
- AI visibility: how often and how well your brand appears in AI answers.
Alvaro Peña de Luna