AI crawler index
What is CCBot?
CCBot is Common Crawl's crawler for AI training. Builds Common Crawl's free, open repository of web crawl data, a common source for training datasets.
At a glance
| Operator | Common Crawl |
|---|---|
| Purpose | AI training |
| User agent | CCBot/2.0 (https://commoncrawl.org/faq/) |
| Obeys robots.txt | Yes, according to the operator |
| Published IP ranges | https://index.commoncrawl.org/ccbot.json |
| How to verify it | Reverse DNS to *.crawl.commoncrawl.org and forward-confirm (IPv4 only), or match the IP list. |
| Official documentation | https://commoncrawl.org/ccbot |
Should you block CCBot?
It collects pages for model training, not for answering questions in real time. Blocking it keeps your future content out of that operator's training data, but it does not take you out of AI search answers, which other crawlers feed. Block it if you do not want your content in training sets; allow it if you want future models to know your brand.
It obeys Crawl-delay. Blocking it stops future crawls only, not the archives already published. Common Crawl warns that fake CCBots exist, so verify it.
How to block or allow CCBot in robots.txt
Block it everywhere:
User-agent: CCBot
Disallow: / Allow it everywhere:
User-agent: CCBot
Allow: / Rules are matched by the token on the User-agent line. Test the result with the robots.txt tester before you publish it.
Check your own site
The AI crawler checker tells you which AI crawlers your robots.txt lets in, and the log file analyzer shows how often each one actually visits. For the bigger picture, see what AI crawlers are and the full AI crawler index.
FAQ
CCBot questions
Start tracking your brand in AI search today
Monitor how AI engines cite your brand, track keyword positions, and benchmark against competitors, all in one platform.