Retrieval-augmented generation (RAG) is a technique in which an AI system first retrieves relevant documents from an external source, such as a search index or a database, and then passes them to a large language model so it writes its answer from that material instead of relying only on what it learned in training.
RAG is the reason AI search engines can talk about last week’s news, show sources under their answers and cite your pages. For brands, it is the part of AI search you can influence most directly.
How retrieval-augmented generation works
The term comes from a 2020 research paper by researchers at Facebook AI Research (now Meta). The idea is simple: a large language model is good at writing but its memory is frozen and imprecise, so you give it the right documents at the moment it answers. In an AI search engine, the process usually looks like this:
- Understand the question. The system rewrites the prompt into one or more search queries, sometimes many at once (see query fan-out).
- Retrieve. It runs those queries against an index. Retrieval can match keywords, meaning (using vector embeddings) or both.
- Select passages. From the results, it keeps the most relevant fragments, often a few paragraphs per page rather than whole pages.
- Augment the prompt. Those passages are inserted into the instructions the model receives, together with the user’s question.
- Generate. The model writes the answer from the passages and, in most AI search engines, attaches the pages it used as sources.
RAG in AI search engines
Every major AI search experience uses some form of retrieval. Perplexity is built around searching the web and citing sources. ChatGPT can run a web search when it judges the question needs one. Google AI Overviews and Google AI Mode draw on Google’s own search index. The details differ, but the pattern is the same: the answer is only as good, and only as favorable to you, as the documents that were retrieved.
Why RAG matters for your brand
- It moves the competition to retrieval. You cannot edit a model’s training data, but you can publish and earn pages that are retrieved for the questions your customers ask.
- Passages compete, not pages. Retrieval often works on fragments. A clear, self-contained section that answers one question is easier to select than a long page that mixes several topics.
- Third parties speak for you. If the retrieved sources are review sites, comparisons and articles, their description of your brand is what the model repeats.
- Freshness counts. A retrieved page reflects its current content, so updated pricing or a new product can reach answers without waiting for a new model.
How to optimize for retrieval
- Be crawlable. Check that search and AI crawlers can reach your key pages with the robots.txt tester.
- Answer in self-contained passages. Use descriptive headings and put the answer in the first sentences below them.
- Keep facts in text. Prices, features and specifications locked in images or scripts are harder to retrieve than plain HTML.
- Rank in classic search. Engines that retrieve from a search index tend to draw on pages that already perform there.
- Be present in the sources that get cited. Find which third-party pages appear in answers for your category and work on being included in them.
Common mistakes
- Confusing a citation with a mention. An engine can cite your page without naming your brand, or name you while citing someone else. Track both.
- Assuming the engine reads your whole site. It reads the few passages it retrieved for that question, nothing more.
- Assuming retrieval always happens. Some answers come from the model’s memory alone, with no sources at all.
How Mencoro shows what AI engines retrieve
Mencoro records the text of each AI answer and the sources it returns. ChatGPT is asked with web search on, and its sources include the pages its search retrieved even when the text does not cite them. A link counts as yours when it points to one of your domains or subdomains, even if the answer never names your brand, and its Link position is its place in the list of sources. Links feed Link position and Share of voice (a link to your domain counts as a direct, neutral reference with a weight of 0.14), but not Coverage or Favorability, which are about what the answer says. The full method is on how Mencoro works, and the ChatGPT rank tracker shows it applied to ChatGPT.
Related glossary terms
- Grounding: anchoring an AI answer in verifiable sources.
- Vector embeddings: the numeric representations that let retrieval match meaning.
- AI citations: the sources an AI answer links to.
- Content chunking: splitting content into passages that can be retrieved on their own.
Alvaro Peña de Luna