A knowledge cutoff (or training data cutoff) is the date after which a large language model has no information from its training data: anything published or changed after that point is unknown to the model unless it is retrieved from the web or another source at the moment it answers.
For a brand, the cutoff explains a familiar frustration: an AI answer that describes your old pricing, ignores your new product or does not know you exist.
Why language models have a knowledge cutoff
A large language model learns from a snapshot of text collected up to a certain date. Training and testing the model then take months, so the cutoff is usually earlier than the release date, and the model stays in use long after that. Model providers often state the cutoff in their documentation.
Two details are easy to miss:
- The last months are thin. Less has been written about recent events when the snapshot is taken, so the model tends to know less about the period just before its cutoff.
- Models are unreliable about their own cutoff. Asking a chatbot for its cutoff date is not a dependable way to find it.
Knowledge cutoff vs web search
AI search engines work around the cutoff with retrieval-augmented generation: they search the web and give the results to the model, which is how they can answer about recent events. But the frozen memory does not go away. It is still there in every answer, and it takes over whenever the engine does not search. In ChatGPT, for example, the model decides whether a question needs a web search. When it does not search, it answers from training data, and the answer describes your brand as it was before the cutoff.
What the knowledge cutoff means for your brand
- If you launched after the cutoff, you are invisible in answers from memory. Your only way into AI answers for now is retrieval, which depends on your pages and your coverage being found.
- If you rebranded, repriced or changed products, the model’s memory is stale. Answers can mix old and new facts, or repeat the old ones as hallucinations.
- If you are an established brand, training data gives you default associations with your category. Losing ground in current coverage shows up first in grounded answers, long before a new model reflects it.
This is why fresh coverage matters twice. Today, recent reviews, comparisons, press and updated pages are what search-enabled engines retrieve and cite. Tomorrow, that same coverage is part of the text future models are trained on.
How to work with the knowledge cutoff
- Keep key facts current and in plain text on your site: pricing, products, locations and what changed.
- Announce changes where they will be found. A crawlable page on your site plus coverage in third-party sources that already rank for your category.
- Update old third-party pages. Directories, comparisons and reviews with outdated facts feed both retrieval and future training.
- Redirect and explain old names. If you renamed a product, say so on the page that replaces it.
- Stay crawlable. Check your rules with the robots.txt tester so search and AI crawlers can reach the pages that carry the new facts.
Common mistakes
- Assuming a search-enabled engine always searches. It often answers from memory.
- Expecting instant updates. A change on your site reaches grounded answers only after it is crawled and retrieved, and answers from memory only with a new model.
- Only updating your own site. The model learned about you mostly from what others wrote.
How Mencoro helps you see the cutoff at work
Mencoro asks ChatGPT with web search on, but the model decides whether to search: when it answers from what it already knows, the answer has no sources, and the mentions in its text are still counted. Reading those answers next to the ones with sources shows which descriptions of your brand come from memory and which from pages retrieved today. Mencoro keeps 16 months of history, so you can see when a new product or price starts to appear in answers on ChatGPT, Perplexity, Google AI Overview and Google AI Mode. The method is on how Mencoro works, and the ChatGPT rank tracker and AI visibility feature show it in practice.
Related glossary terms
- Large language model (LLM): the type of model that has a knowledge cutoff.
- Grounding: how engines tie answers to current sources.
- AI hallucination: false claims, often caused by outdated knowledge.
- Retrieval-augmented generation (RAG): the technique that brings fresh documents into an answer.
Alvaro Peña de Luna