Retrieval-Augmented Generation (RAG)
RAG is a pattern that retrieves relevant information from an external knowledge source and provides it to a language model before generating an answer.
Prerequisites
Why RAG?
A language model's knowledge comes entirely from what it saw during training. That creates real limitations once you try to build something useful with it.
- Knowledge cutoff — the model has no information about anything that happened after its training data was collected.
- Private data — a model has never seen your company's internal documents, policies, or product data, because that data was never public.
- Changing information — prices, policies, and product details change constantly; a model's internal knowledge is frozen at training time.
- Hallucination risk — when a model doesn't actually know something, it can still generate a fluent, confident-sounding answer that is wrong.
- Domain-specific knowledge — highly specialized or niche information is often underrepresented in general training data.
RAG doesn't try to make the model smarter or retrain it. It changes what information the model has available at the moment it answers, by retrieving relevant content and giving it to the model as context.
Key Idea
RAG grounds generation in retrieved information instead of relying only on what the model memorized during training.
A Simple Explanation
Without RAG, a model can only work with what it already knows. RAG adds one extra step before the model ever generates anything:
Question
LLM
Answer
Question
Find relevant information
Give information to LLM
Generate answer
That extra step is the whole idea: before the model generates anything, the system searches for information relevant to the question and hands it to the model as part of the prompt.
The Mental Model
Key Idea
RAG = retrieve useful context, then generate the answer.
Four ideas do most of the work here, and they map directly onto that sentence:
- Query — the question or request you want answered.
- Retrieval — searching a knowledge source for information relevant to that query.
- Context — the retrieved information, assembled into something the model can read alongside the question.
- Generation — the model producing an answer using the question and the retrieved context together.
Hold onto that before looking at implementation details: a retrieval step happens first, and its output becomes part of what the model reads before it answers.
How RAG Works
At the moment a user asks a question, a RAG system runs a query-time pipeline:
User Question
Starting PointWhat the user actually wants to know.
Query Processing
Prepares for SearchCleans up or rewrites the raw question.
Retriever
Finds CandidatesSearches the prepared knowledge source.
Vector / Keyword / Hybrid Search
Search MethodHow the retriever actually matches content.
Relevant Documents
Raw MatchesWhat the search step found.
Context Assembly
Builds the PromptCombines documents into what the model reads.
LLM
Grounded GenerationAnswers using the assembled context.
Generated Answer
Final OutputGrounded in retrieved documents, not memory alone.
That retrieval step only works because the knowledge source was prepared in advance. Before any question is ever asked, documents go through an ingestion pipeline:
Documents
Raw SourceThe knowledge base before any processing.
Parsing
Extracts TextTurns raw files into usable text.
Chunking
Splits ContentSmaller pieces sized for retrieval.
Embedding
Vector FormEach chunk converted for similarity search.
Vector Database
Ready to SearchWhat the retriever queries at question time.
Ingestion happens once, and again whenever documents change. Retrieval happens every single time a user asks a question.
Key Components
Here's how the pieces from the diagrams above relate to each other:
- Documents
- The raw source material — PDFs, wiki pages, tickets, policies — that contains the knowledge you want the system to use.
- Chunking
- Splitting documents into smaller pieces so they can be embedded and retrieved effectively, instead of retrieving an entire document at once.
- Embeddings
- Numerical vector representations of text that let the system compare meaning, not just exact words.
- Vector Database
- Storage optimized for finding the embeddings most similar to a query's embedding, quickly, at scale.
- Retrieval
- The step of actually searching the knowledge source and returning the most relevant chunks.
- Reranking
- An optional second pass that reorders retrieved results using a more precise (and more expensive) model.
- Context
- The retrieved chunks, assembled into the prompt the model will actually read.
- LLM
- The model that reads the question and the retrieved context, and generates the final answer.
A Real-World Example
A company has thousands of internal documents — policies, handbooks, and process guides. An employee asks an internal assistant:
"What is our parental leave policy?"
A RAG system handling this:
- Receives the question.
- Searches the internal knowledge base for relevant content.
- Retrieves the sections of the employee handbook that mention parental leave.
- Sends those retrieved sections to the LLM as context, along with the original question.
- Generates an answer grounded in the retrieved policy text.
Important
The model is not querying the company database by itself. Retrieval is a separate, explicit step that happens before generation — the model only ever sees the text it was handed.
Code Example
This is a minimal illustration of the pattern, not a production implementation:
def answer_question(question, retriever, llm):
# 1. Retrieve documents relevant to the question
documents = retriever.search(question, top_k=3)
# 2. Build context from the retrieved documents
context = "\n\n".join(doc.text for doc in documents)
# 3. Send the question and context to the model together
prompt = f"""Answer the question using only the context below.
Context:
{context}
Question: {question}
Answer:"""
return llm.generate(prompt)Real systems add a chunking strategy, an embedding model, a vector database, metadata filtering, and often a reranking step — but the core pattern stays the same: retrieve, then generate.
Common Mistakes
Poor chunking
Chunks that are too large dilute relevance; chunks that are too small lose context. Both hurt retrieval quality.
Retrieving irrelevant documents
If the retriever returns weakly related content, the model will generate an answer grounded in the wrong information.
Retrieving too much context
Stuffing the prompt with excessive retrieved text increases cost and can bury the passage that actually matters.
Ignoring metadata
Filtering by source, date, or permissions is often as important as semantic similarity for returning the right content.
Assuming vector search is always enough
Keyword and hybrid search often outperform pure vector search for exact terms, IDs, or rare vocabulary.
Skipping evaluation
Without measuring retrieval quality and answer accuracy, it's hard to know whether changes actually improve the system.
Confusing retrieval problems with LLM problems
A wrong answer is often caused by bad retrieval, not a weak model — the model can only work with what it was given.
Failing to inspect retrieved context
The fastest way to debug a bad RAG answer is to look at what was actually retrieved, not just the final output.
Interview Question
What is RAG and why would you use it?
RAG, or Retrieval-Augmented Generation, retrieves relevant information from an external knowledge source and provides it to a language model as context before it generates a response. You'd use it when an application needs to answer using information the model wasn't trained on — private company data, frequently changing information, or domain-specific knowledge — without retraining or fine-tuning the model itself. It also reduces hallucination risk by grounding the answer in retrieved, verifiable content instead of relying purely on what the model memorized.
What an interviewer may ask next
- Why not just fine-tune the model instead?
- What role do embeddings play in retrieval?
- What is chunking, and why does it matter?
- Why would you add a reranking step?
- How would you evaluate whether a RAG system is actually working well?
Explain It in 30 Seconds
RAG is a pattern where you retrieve relevant external information and give it to a language model as context before it generates a response. It lets an application answer using private, current, or domain-specific knowledge without requiring the model itself to memorize that information — and because the answer is grounded in retrieved content, it's less likely to be a confident-sounding guess.