AI crawler index

AI crawler index

The crawlers AI companies send to your site, what each one is for and whether it obeys robots.txt. Every fact comes from the operator's own documentation, checked on 1 October 2026.

All AI crawlers

CrawlerOperatorPurposeObeys robots.txt
Claude-SearchBot Anthropic AI search Yes
DuckAssistBot DuckDuckGo AI search Yes
OAI-SearchBot OpenAI AI search Yes
PerplexityBot Perplexity AI search Yes
YouBot You.com AI search Yes
ChatGPT-User OpenAI User-triggered fetch Not always
Claude-User Anthropic User-triggered fetch Yes
meta-externalfetcher Meta User-triggered fetch Not always
MistralAI-User Mistral AI User-triggered fetch Yes
Perplexity-User Perplexity User-triggered fetch Not always
Amazonbot Amazon AI training Yes
CCBot Common Crawl AI training Yes
ClaudeBot Anthropic AI training Yes
GPTBot OpenAI AI training Yes
meta-externalagent Meta AI training Yes
Applebot-Extended Apple robots.txt control token Yes
Google-Extended Google robots.txt control token Yes
Applebot Apple Search engine Yes
Bingbot Microsoft Search engine Yes
Googlebot Google Search engine Yes
PetalBot Huawei Search engine Yes
Bytespider ByteDance Other or undeclared Not documented
GoogleOther Google Other or undeclared Yes

Four kinds of AI crawler

  • AI search crawlers build the index an assistant cites from (OAI-SearchBot, Claude-SearchBot, PerplexityBot). Block them and you can disappear from those answers.
  • User-triggered fetchers read a page only when a user asks (ChatGPT-User, Claude-User, Perplexity-User). Some operators say robots.txt may not apply to them.
  • Training crawlers collect pages that may train models (GPTBot, ClaudeBot, CCBot). Blocking them keeps your future content out of training, not out of search.
  • Control tokens have no crawler of their own (Google-Extended, Applebot-Extended): they decide how the main search crawler's data is used for AI.

robots.txt templates

Stay in AI search, stay out of AI training

Allows the crawlers that put you in AI answers and blocks the ones that only collect training data. This is the setup most brands that want AI visibility choose.

User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-externalfetcher
User-agent: DuckAssistBot
User-agent: MistralAI-User
User-agent: YouBot
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Amazonbot
User-agent: meta-externalagent
User-agent: CCBot
User-agent: Bytespider
Disallow: /

Block every AI crawler

Keeps AI crawlers out but leaves search engines such as Googlebot and Bingbot alone. Expect to lose citations in ChatGPT, Claude and Perplexity answers.

User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-externalfetcher
User-agent: DuckAssistBot
User-agent: MistralAI-User
User-agent: YouBot
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Amazonbot
User-agent: meta-externalagent
User-agent: CCBot
User-agent: Bytespider
Disallow: /

Paste the template into your robots.txt, then test it with the robots.txt tester. User-triggered fetchers that ignore robots.txt need a firewall rule as well.

Which crawlers can reach your site?

The free AI crawler checker reads your robots.txt and requests your page as each crawler. The log file analyzer shows which ones actually visit and what they ask for. New to the topic? Start with what AI crawlers are.

Start tracking your brand in AI search today

Monitor how AI engines cite your brand, track keyword positions, and benchmark against competitors, all in one platform.