Free tool

AI crawler checker: can ChatGPT, Claude and Perplexity read your site?

Check in seconds whether 16 AI crawlers can read a page, from your robots.txt rules to how your server answers them.

What this AI crawler checker tests

An AI engine can only recommend what it can read. This AI crawler checker looks at the four places where a site can shut AI crawlers out, often without anyone noticing:

  1. robots.txt. Each crawler’s rules are applied the way the crawler reads them: the most specific user-agent group wins and the longest matching path decides.
  2. Your server and CDN. We request the page with the user agent of each main crawler and compare the answer with a normal browser. Firewalls that block “AI bots” with one click show up here as a 403.
  3. Page directives. A robots meta tag or an X-Robots-Tag header with noindex or nosnippet keeps a page out of search features, AI Overviews included.
  4. llms.txt. Whether the site publishes one. It is optional, and our llms.txt generator drafts one from your sitemap.

The AI crawlers we check

Crawler Company Used for
GPTBot OpenAI Model training
OAI-SearchBot OpenAI AI search index
ChatGPT-User OpenAI Visits when a user asks
ClaudeBot Anthropic Model training
Claude-SearchBot Anthropic AI search index
Claude-User Anthropic Visits when a user asks
PerplexityBot Perplexity AI search index
Perplexity-User Perplexity Visits when a user asks
Googlebot Google AI search index
Google-Extended Google Opt-out token, no crawler of its own
Bingbot Microsoft AI search index
Applebot-Extended Apple Opt-out token, no crawler of its own
meta-externalagent Meta Model training
Amazonbot Amazon AI search index
CCBot Common Crawl Model training
Bytespider ByteDance Model training

The split that matters is between training and search. Blocking a training crawler keeps your content out of future models; blocking a search crawler removes you from the answers that engine gives today.

How to allow or block AI crawlers in robots.txt

To stay visible in AI search while opting out of model training, allow the search and user-triggered crawlers and disallow the training ones:

User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /

Two things to remember. robots.txt is a request, not a lock, so check your CDN settings as well, since that is where most accidental blocks live. And every change here trades training exposure against visibility: decide per crawler, not with one global rule. To see which of them actually visit you, and how often, drop your access log into the log file analyzer . To try a rule before you publish it, paste your edited file into the robots.txt tester and test any URL against it.

Why crawler access matters for AI visibility

Crawler access is the floor, not the goal. A site that AI engines can read can still be left out of the answers, and the only way to know is to measure. Mencoro tracks how ChatGPT, Perplexity and Google’s AI features mention your brand , next to your competitors, so you can see whether opening the door actually changes the answer.

FAQ

AI crawler checker: frequently asked questions

Enter your domain above. The checker reads your robots.txt and applies each AI crawler’s rules the way the crawler does, then requests the page with the user agents of the main crawlers to see whether your server or firewall answers them differently from a normal browser.
It depends on what you want. GPTBot collects data to train OpenAI’s models; OAI-SearchBot builds the index that ChatGPT search cites. Blocking GPTBot keeps your content out of training without removing you from ChatGPT search, as long as OAI-SearchBot stays allowed.
No. Google-Extended controls whether your content is used to train and ground Gemini models. AI Overviews and AI Mode use Googlebot’s regular crawl, so the way to stay out of them is a snippet directive such as nosnippet, not Google-Extended.
Firewalls and CDNs such as Cloudflare can block AI bots by user agent regardless of robots.txt. If the checker shows 403 for a crawler your robots.txt allows, the block is in your server or CDN settings. For Googlebot and Bingbot, a 403 can also mean your firewall blocks impostors, which is fine.

Start tracking your brand in AI search today

Monitor how AI engines cite your brand, track keyword positions, and benchmark against competitors, all in one platform.