GPTBot is OpenAI's web crawler for collecting publicly available pages that may be used to train its generative AI models. It identifies itself with the GPTBot user agent token, reads robots.txt, and is separate from OAI-SearchBot, the crawler behind ChatGPT search.
That separation is the one thing to remember about GPTBot: it decides what OpenAI's future models may learn from your site, not whether ChatGPT can cite you today.
How GPTBot works
GPTBot requests public pages like any crawler, reads their content and follows links. Its full
user agent string contains GPTBot and a link to openai.com/gptbot, and
OpenAI publishes the IP ranges its crawlers use, so you can tell a real visit from an impostor.
OpenAI documented GPTBot in August 2023, and its current description lives on the OpenAI crawlers page.
What GPTBot collects does not appear in ChatGPT the next day. It may become part of the training data of a future model, which only reflects it after that model is trained and released, and only up to its knowledge cutoff.
GPTBot vs OAI-SearchBot vs ChatGPT-User
OpenAI runs three user agents, and each one can be allowed or blocked on its own in robots.txt:
| User agent | What it does | If you block it |
|---|---|---|
GPTBot | Collects content that may be used to train OpenAI's models | Your new content stays out of future training data |
OAI-SearchBot | Builds the index ChatGPT search uses to find and cite pages | Your pages stop being retrieved and cited in ChatGPT search answers |
ChatGPT-User | Fetches a page when a user asks ChatGPT to read or check it | ChatGPT cannot open your pages on a user's request |
Other AI companies follow a similar split, which the AI crawlers entry covers vendor by vendor.
Should you block GPTBot?
It is a business decision, not a technical one. Publishers whose content is the product often block GPTBot to keep it out of training. Brands that want AI models to know who they are, what they sell and how they compare usually allow it, because training data is part of what a model draws on when it answers without searching.
Whatever you decide about GPTBot, the crawler that affects your visibility in ChatGPT answers today is OAI-SearchBot. Blocking GPTBot while allowing OAI-SearchBot is a common, consistent policy.
How to block or allow GPTBot in robots.txt
To block training while staying visible in ChatGPT search:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Allow: / To let GPTBot read only one section, such as your blog:
User-agent: GPTBot
Allow: /blog/
Disallow: /
The second example works because the longest matching rule wins: /blog/ is longer than
/, so blog URLs are allowed and everything else is blocked. Keep in mind that once
GPTBot has its own group, it ignores your User-agent: * rules, and that a site with no
GPTBot group and a general Disallow: / blocks GPTBot too. Test your edited file in the
free robots.txt tester before you publish it.
How to check whether GPTBot visits your site
- Access: the AI crawler checker applies your robots.txt to GPTBot, OAI-SearchBot and ChatGPT-User and requests your page with their user agents, so a firewall or CDN block shows up as a 403.
- Visits: the log file analyzer reads your access log in your browser and shows how often each OpenAI crawler comes, which pages it reads and where it gets errors.
- Authenticity: compare the IP addresses of suspicious requests with the ranges OpenAI publishes. A request that says GPTBot from an unlisted IP is not OpenAI.
Common mistakes
- Blocking GPTBot to leave ChatGPT search. It does not work. Search depends on OAI-SearchBot.
- Blocking every OpenAI bot in one go. A CDN rule against "AI bots" usually catches OAI-SearchBot and ChatGPT-User as well, which removes your pages from ChatGPT's sources.
- Expecting instant effects. Allowing GPTBot does not change today's answers, and blocking it does not erase what models already learned.
Measuring the effect on ChatGPT answers
Crawler rules only matter if they change what ChatGPT says about you. Mencoro asks ChatGPT your prompts with web search on and records both the text of the answer and the sources it returns, including the pages its search retrieved. The ChatGPT rank tracker shows how often you are named, in what position and whether your own pages are among the sources, next to your competitors. For a one-off sample, try the free ChatGPT rank checker.
For every AI crawler with its user agent, IP ranges and robots.txt behaviour, see the AI crawler index.
Related glossary terms
- AI crawlers: every AI bot, grouped by what it collects for.
- Knowledge cutoff: the date after which a model's training data stops.
- AI referral traffic: visits from people who click links in AI answers.
- llms.txt: a curated map of your key pages for AI models.
Alvaro Peña de Luna