Retrieval-Augmented Generation: How an AI Model Uses External Knowledge
AI Foundations #23 explains retrieval-augmented generation, embeddings, chunking, retrieval, context injection, citations, failure modes, and why RAG is different from training.
Approximately 12 min read · AI Foundations / Lesson 23
AI Foundations #22 asked how we know whether a model is better.
That lesson focused on evaluation: test sets, metrics, benchmark design, variance, and contamination.
Now we can ask a different question:
What should a model do when the answer is not safely stored in its parameters?
A language model can know a great deal from training, but its internal weights are not a live database.
The model may not contain a recent fact.
It may not know a private document.
It may remember a fact imperfectly.
It may have learned two conflicting versions of the same information.
One answer is retrieval-augmented generation, usually shortened to RAG.
RAG separates two jobs:
- find relevant information;
- generate an answer using that information.
That sounds simple, but the separation changes how we think about knowledge in an AI system.
The basic idea
A plain language-model request looks roughly like this:
question
|
v
language model
|
v
answer
A retrieval-augmented request adds an information-retrieval stage:
question
|
v
retriever
|
v
relevant passages
|
+------+
|
v
question + passages
|
v
language model
|
v
answer
The model still generates text token by token.
The difference is that some of the evidence can be supplied at inference time instead of being recovered only from model parameters.
Why not just train the fact into the model?
Training and retrieval solve different problems.
Training changes model parameters.
Retrieval usually leaves model parameters unchanged and changes the context supplied for this request.
That distinction matters.
Suppose a company changes an internal support policy this morning.
Retraining or fine-tuning a large model for one policy update would be expensive and slow.
A retrieval system can instead update the document collection, retrieve the new policy when relevant, and place it into the prompt.
So:
training
-> changes the model
retrieval
-> changes the information shown to the model
RAG is not a replacement for good training.
It is a way to connect a trained model to information that can be updated, selected, or kept outside the model.
Step 1: prepare the knowledge source
A RAG system starts with material that may be useful later:
- manuals
- policies
- research papers
- support articles
- product documentation
- database records
- notes
- source code
- transcripts
Large documents are usually divided into smaller pieces called chunks.
Why?
Because retrieving an entire 300-page manual for every question would waste context and make it harder to isolate the useful section.
A simple document pipeline might look like:
document
|
v
split into chunks
|
v
represent each chunk
|
v
store searchable index
Chunking is a design choice.
Chunks that are too small can lose surrounding meaning.
Chunks that are too large can contain lots of irrelevant text.
There is no universal perfect chunk size.
Step 2: represent meaning
One common RAG design uses embeddings.
Earlier in AI Foundations #9, we saw that an embedding is a vector representation.
For retrieval, the rough goal is:
passages with similar meaning should have vectors that are close under some similarity measure.
Imagine three chunks:
A: "The server uses 24 GB of VRAM."
B: "The GPu memory requirement is twenty-four gigabytes."
C: "The office closes at 6 PM."
A useful embedding model should place A and B closer together than A and C, even though A and B use different words.
The user’s query is embedded too.
Then the retrieval system searches for nearby document vectors.
Step 3: retrieve candidate passages
Suppose the user asks:
How much GPU memory does the server need?
The system can convert that question into a query representation and retrieve the most relevant chunks.
In a simplified vector-search system:
query embedding
|
v
compare with stored chunk embeddings
|
v
top matching chunks
The retriever is not generating the final answer.
It is selecting evidence for the generator.
This division of labor is central to RAG.
The original Dense Passage Retrieval work showed how learned dense representations could retrieve passages for open-domain question answering.
The original RAG paper then combined retrieval with sequence generation so retrieved documents could support generated answers.
Retrieval does not have to be vector search
“RAG” is often casually used as if it means “vector database.”
That is too narrow.
Retrieval can use:
- keyword search
- BM25-style lexical ranking
- dense vector search
- sparse learned retrieval
- database queries
- metadata filters
- graph traversal
- hybrid retrieval
- rerankers
- combinations of these
A system might first filter documents by product version, then run keyword and vector search, then rerank the candidates.
The important property is not a particular database.
The important property is that external information is selected and supplied to generation.
Step 4: put evidence into the model context
After retrieval, selected passages are inserted into the model’s input.
For example:
SYSTEM:
Answer using the supplied documentation.
CONTEXT:
[passage 1]
The server requires 24 GB of GPU memory...
[passage 2]
For long-context mode, reserve additional memory...
USER:
How much GPU memory does the server need?
The language model now receives evidence that was not required to live entirely in its weights.
This reconnects RAG to context length.
Retrieved text consumes context tokens.
More retrieved text is not automatically better.
If the system retrieves too much material, useful evidence can be surrounded by noise.
Retrieval quality and generation quality are different
A RAG answer can fail in at least two broad places.
Retrieval failure
The right document exists, but the system does not retrieve it.
Then the generator never sees the evidence.
Generation failure
The right evidence is retrieved, but the model misreads it, ignores it, combines it incorrectly, or adds unsupported claims.
These failures require different debugging.
If the retriever returned the wrong passages, changing the answer prompt alone may not solve the real problem.
If retrieval is excellent but the model keeps inventing unsupported details, the retrieval layer may not be the bottleneck.
This is why RAG evaluation often needs more than one score.
A useful mental model
Think of a student taking an exam.
A closed-book exam tests what the student can recall from memory.
An open-book exam gives the student access to references.
But giving someone a book does not guarantee a correct answer.
They still need to:
- search for the right section;
- recognize whether it is relevant; . interpret it correctly; . answer the actual question.
RAG gives the model something closer to an open-book workflow.
The retriever finds pages.
The generator reads the selected pages and produces the response.
RAG can improve freshness
Model parameters are tied to a training process.
A retrieval index can often be updated much faster.
That is useful for information such as:
- changing product documentation
- internal procedures
- inventories
- current research collections
- frequently updated knowledge bases
But freshness depends on the retrieval source actually being current.
An old index does not become fresh just because it is connected to a language model.
RAG can provide private or domain-specific context
A general model may never have seen an organization’s private documents during training.
A retrieval system can make authorized internal material available at inference time.
The model does not need to have memorized every company-specific detail during pretraining.
This is one reason RAG became important in enterprise AI systems.
But access control remains essential.
A retriever must not return documents the user is not authorized to see.
RAG does not automatically solve permissions.
RAG and hallucination
RAG is often described as a way to reduce hallucination.
That can be true when the system retrieves strong evidence and the model follows it.
But retrieval does not mathematically force the model to stay faithful to the evidence.
The model can still:
- misunderstand a passage
- combine unrelated passages
- answer beyond the evidence
- prefer prior knowledge over retrieved text
- cite a passage that does not support the claim
So a better statement is:
RAG can give the model relevant external evidence and make grounded answers easier, but grounding still has to be designed and evaluated.
Citations are another layer
A RAG system may show citations with its answer.
That is useful only if the cited source actually supports the nearby claim.
A weak system might retrieve five documents, generate an answer, and attach all five links at the bottom.
That is not the same as claim-level evidence.
A stronger workflow tracks which retrieved passage supports which part of the answer.
The citation mechanism is separate from retrieval itself.
Top-k retrieval
Many systems retrieve the top k candidates.
If k = 5, the model may receive five passages.
Increasing k can improve recall because there are more chances to include the right evidence.
But increasing k can also add noise and consume context.
This creates a tradeoff:
small k
-> less noise
-> lower chance of including the needed passage
large k
-> better chance of including the needed passage
-> more noise and context use
The best value depends on the documents, retriever, reranker, model, and task.
Reranking
A common pipeline retrieves a larger candidate set cheaply, then uses a stronger model to rerank it.
For example:
10,000,000 chunks
|
v
fast retriever
|
v
top 50
|
v
reranker
|
v
top 5
|
v
language model
This separates broad recall from precise ranking.
The first stage tries not to miss useful material.
The reranker spends more computation deciding which candidates are most relevant.
Metadata filters
Meaning alone is not always enough.
Suppose a knowledge base contains documentation for software versions 2, 3, and 4.
The question concerns version 4.
A semantically similar version-2 passage may be dangerous.
Metadata can narrow the search:
product = X
version = 4
language = English
document_status = current
Then semantic retrieval operates within the allowed subset.
This is especially important when old and new documentation look very similar.
Hybrid retrieval
Keyword and vector search have different strengths.
Keyword search can be excellent for exact identifiers:
CUDA_ERROR_OUT_OF_MEMORY
PR #58890
Qwen3-Omni
Dense retrieval can help when the query and document express the same idea with different wording.
Hybrid systems combine these signals rather than pretending one method is always best.
RAG is not fine-tuning
These are commonly confused.
Fine-tuning changes parameters.
RAG supplies external context.
A rough comparison:
| Question | Fine-tuning | RAG |
|---|---|---|
| Changes model weights? | Yes | Usually no |
| Easy to update one document? | No | Yes |
| Good for teaching response style? | Often | Not its main purpose |
| Good for current external facts? | Limited | Often |
| Requires retrieval infrastructure? | No | Yes |
Some systems use both.
Fine-tuning can shape behavior while RAG supplies changing knowledge.
RAG is not a guarantee of truth
A retrieval system can retrieve bad sources.
The source itself can be wrong.
The index can be stale.
The query can be ambiguous.
The ranking can be poor.
The model can interpret evidence incorrectly.
So the system’s reliability depends on the whole chain:
source quality
-> indexing
-> retrieval
-> ranking
-> context construction
-> generation
-> citation / verification
Improving only the final language model may leave the real failure untouched.
How to evaluate a RAG system
AI Foundations #22 introduced the idea that evaluation must match the system.
For RAG, useful questions include:
Did retrieval find the needed evidence?
This tests the retrieval stage.
Was the correct passage ranked high enough?
Finding the right chunk at rank 200 may not help if only the top five reach the model.
Did the answer use the evidence correctly?
This tests generation and grounding.
Are claims supported by the cited passages?
This tests attribution.
Does the system abstain when evidence is missing?
Sometimes the correct behavior is to say that the provided sources do not answer the question.
A single overall accuracy number can hide where the pipeline failed.
The original RAG idea
The 2020 RAG paper described a system combining a pretrained sequence-to-sequence model with a non-parametric memory accessed through dense retrieval.
The important conceptual split still matters today:
- parametric memory: information represented in model parameters;
- non-parametric memory: external information retrieved at inference time.
Modern production systems may look very different from the exact research architecture in that paper.
They may use different retrievers, rerankers, vector databases, prompt formats, or models.
But the core idea remains recognizable:
retrieve evidence, then condition generation on it.
REALM and retrieval during language modeling
REALM explored retrieval-augmented language-model pretraining, showing another important direction: retrieval can be integrated more deeply than simply adding a search step after a model is trained.
This reminds us that “RAG” describes a family of architectures, not one fixed implementation.
Some systems retrieve only at application time.
Others train models to interact with retrieval.
Where RAG fits in the Foundations map
We can now connect several earlier lessons:
tokens
-> embeddings
-> attention / Transformers
-> context
-> inference
-> evaluation
-> retrieval-augmented generation
Embeddings can represent queries and passages.
Context carries retrieved evidence into the model.
Inference generates the answer.
Evaluation tells us whether retrieval and generation are actually working.
RAG is therefore not an isolated trick.
It combines ideas from many parts of the AI stack.
What RAG does not solve
RAG does not automatically solve:
- reasoning
- source quality
- authorization
- contradictory documents
- malicious retrieved text
- prompt injection
- long-context limits
- poor chunking
- bad ranking
- unsupported generation
- citation accuracy
These become system-design problems.
That is why a real RAG application is more than “put PDFs in a vector database.”
The most important distinction
The cleanest way to remember RAG is:
The model does not have to know everything before the request begins.
It can retrieve external information, place that information into its working context, and generate from both its learned parameters and the retrieved evidence.
That gives AI systems a practical path to fresher, private, or domain-specific knowledge.
But retrieval adds a new chain of possible failures, so it must be evaluated as a system rather than treated as an automatic truth layer.
The next Foundations lesson can build on this by asking what happens when a model does more than retrieve information and begins choosing tools and actions.