What is GPTBot? OpenAI's crawler and how to allow or block it

GPTBot is OpenAI's web crawler for model training. How it differs from OAI-SearchBot and ChatGPT-User, and how to allow or block it in robots.txt.

  • Alvaro Peña de Luna Alvaro Peña de Luna
  • date icon

    Wednesday, Sep 30, 2026

GPTBot is OpenAI's web crawler for collecting publicly available pages that may be used to train its generative AI models. It identifies itself with the GPTBot user agent token, reads robots.txt, and is separate from OAI-SearchBot, the crawler behind ChatGPT search.

That separation is the one thing to remember about GPTBot: it decides what OpenAI's future models may learn from your site, not whether ChatGPT can cite you today.

How GPTBot works

GPTBot requests public pages like any crawler, reads their content and follows links. Its full user agent string contains GPTBot and a link to openai.com/gptbot, and OpenAI publishes the IP ranges its crawlers use, so you can tell a real visit from an impostor. OpenAI documented GPTBot in August 2023, and its current description lives on the OpenAI crawlers page.

What GPTBot collects does not appear in ChatGPT the next day. It may become part of the training data of a future model, which only reflects it after that model is trained and released, and only up to its knowledge cutoff.

GPTBot vs OAI-SearchBot vs ChatGPT-User

OpenAI runs three user agents, and each one can be allowed or blocked on its own in robots.txt:

User agent What it does If you block it
GPTBot Collects content that may be used to train OpenAI's models Your new content stays out of future training data
OAI-SearchBot Builds the index ChatGPT search uses to find and cite pages Your pages stop being retrieved and cited in ChatGPT search answers
ChatGPT-User Fetches a page when a user asks ChatGPT to read or check it ChatGPT cannot open your pages on a user's request

Other AI companies follow a similar split, which the AI crawlers entry covers vendor by vendor.

Should you block GPTBot?

It is a business decision, not a technical one. Publishers whose content is the product often block GPTBot to keep it out of training. Brands that want AI models to know who they are, what they sell and how they compare usually allow it, because training data is part of what a model draws on when it answers without searching.

Whatever you decide about GPTBot, the crawler that affects your visibility in ChatGPT answers today is OAI-SearchBot. Blocking GPTBot while allowing OAI-SearchBot is a common, consistent policy.

How to block or allow GPTBot in robots.txt

To block training while staying visible in ChatGPT search:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Allow: /

To let GPTBot read only one section, such as your blog:

User-agent: GPTBot
Allow: /blog/
Disallow: /

The second example works because the longest matching rule wins: /blog/ is longer than /, so blog URLs are allowed and everything else is blocked. Keep in mind that once GPTBot has its own group, it ignores your User-agent: * rules, and that a site with no GPTBot group and a general Disallow: / blocks GPTBot too. Test your edited file in the free robots.txt tester before you publish it.

How to check whether GPTBot visits your site

  • Access: the AI crawler checker applies your robots.txt to GPTBot, OAI-SearchBot and ChatGPT-User and requests your page with their user agents, so a firewall or CDN block shows up as a 403.
  • Visits: the log file analyzer reads your access log in your browser and shows how often each OpenAI crawler comes, which pages it reads and where it gets errors.
  • Authenticity: compare the IP addresses of suspicious requests with the ranges OpenAI publishes. A request that says GPTBot from an unlisted IP is not OpenAI.

Common mistakes

  • Blocking GPTBot to leave ChatGPT search. It does not work. Search depends on OAI-SearchBot.
  • Blocking every OpenAI bot in one go. A CDN rule against "AI bots" usually catches OAI-SearchBot and ChatGPT-User as well, which removes your pages from ChatGPT's sources.
  • Expecting instant effects. Allowing GPTBot does not change today's answers, and blocking it does not erase what models already learned.

Measuring the effect on ChatGPT answers

Crawler rules only matter if they change what ChatGPT says about you. Mencoro asks ChatGPT your prompts with web search on and records both the text of the answer and the sources it returns, including the pages its search retrieved. The ChatGPT rank tracker shows how often you are named, in what position and whether your own pages are among the sources, next to your competitors. For a one-off sample, try the free ChatGPT rank checker.

For every AI crawler with its user agent, IP ranges and robots.txt behaviour, see the AI crawler index.

FAQ

Frequently asked questions

No. GPTBot collects content that may be used to train OpenAI's models. ChatGPT search relies on OAI-SearchBot, and pages a user asks ChatGPT to read are fetched by ChatGPT-User. As long as those two stay allowed, blocking GPTBot does not stop ChatGPT from finding and citing your pages.
No. A robots.txt rule only affects what GPTBot collects from the moment it reads the rule. It does not remove content that was already collected or change models that were already trained.
Yes. OpenAI documents that GPTBot reads robots.txt and that each of its crawlers can be allowed or blocked on its own token. Because any bot can claim to be GPTBot, check suspicious traffic against the IP ranges OpenAI publishes.
Search your server access logs for the GPTBot user agent, or load the log into Mencoro's free log file analyzer, which counts GPTBot's requests, the pages it reads and the errors it gets, next to OpenAI's other crawlers.

Start tracking your brand in AI search today

Monitor how AI engines cite your brand, track keyword positions, and benchmark against competitors, all in one platform.