AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Intermediate12 min read

Retrieval-Augmented Generation (RAG)

RAG is a pattern that retrieves relevant information from an external knowledge source and provides it to a language model before generating an answer.

Prerequisites

Why RAG?

A language model's knowledge comes entirely from what it saw during training. That creates real limitations once you try to build something useful with it.

  • Knowledge cutoff — the model has no information about anything that happened after its training data was collected.
  • Private data — a model has never seen your company's internal documents, policies, or product data, because that data was never public.
  • Changing information — prices, policies, and product details change constantly; a model's internal knowledge is frozen at training time.
  • Hallucination risk — when a model doesn't actually know something, it can still generate a fluent, confident-sounding answer that is wrong.
  • Domain-specific knowledge — highly specialized or niche information is often underrepresented in general training data.

RAG doesn't try to make the model smarter or retrain it. It changes what information the model has available at the moment it answers, by retrieving relevant content and giving it to the model as context.

Key Idea

RAG grounds generation in retrieved information instead of relying only on what the model memorized during training.

A Simple Explanation

Without RAG, a model can only work with what it already knows. RAG adds one extra step before the model ever generates anything:

LLM alone

Question

LLM

Answer

RAG

Question

Find relevant information

Give information to LLM

Generate answer

That extra step is the whole idea: before the model generates anything, the system searches for information relevant to the question and hands it to the model as part of the prompt.

The Mental Model

Key Idea

RAG = retrieve useful context, then generate the answer.

Four ideas do most of the work here, and they map directly onto that sentence:

  • Query — the question or request you want answered.
  • Retrieval — searching a knowledge source for information relevant to that query.
  • Context — the retrieved information, assembled into something the model can read alongside the question.
  • Generation — the model producing an answer using the question and the retrieved context together.

Hold onto that before looking at implementation details: a retrieval step happens first, and its output becomes part of what the model reads before it answers.

How RAG Works

At the moment a user asks a question, a RAG system runs a query-time pipeline:

User Question

Starting Point

What the user actually wants to know.

prepared by

Query Processing

Prepares for Search

Cleans up or rewrites the raw question.

sent to

Retriever

Finds Candidates

Searches the prepared knowledge source.

runs

Vector / Keyword / Hybrid Search

Search Method

How the retriever actually matches content.

returns

Relevant Documents

Raw Matches

What the search step found.

combined by

Context Assembly

Builds the Prompt

Combines documents into what the model reads.

sent to

LLM

Grounded Generation

Answers using the assembled context.

produces

Generated Answer

Final Output

Grounded in retrieved documents, not memory alone.

RAG at query time

That retrieval step only works because the knowledge source was prepared in advance. Before any question is ever asked, documents go through an ingestion pipeline:

Documents

Raw Source

The knowledge base before any processing.

processed by

Parsing

Extracts Text

Turns raw files into usable text.

split via

Chunking

Splits Content

Smaller pieces sized for retrieval.

converted via

Embedding

Vector Form

Each chunk converted for similarity search.

stored in

Vector Database

Ready to Search

What the retriever queries at question time.

Ingestion (prepared in advance)

Ingestion happens once, and again whenever documents change. Retrieval happens every single time a user asks a question.

Key Components

Here's how the pieces from the diagrams above relate to each other:

Documents
The raw source material — PDFs, wiki pages, tickets, policies — that contains the knowledge you want the system to use.
Chunking
Splitting documents into smaller pieces so they can be embedded and retrieved effectively, instead of retrieving an entire document at once.
Embeddings
Numerical vector representations of text that let the system compare meaning, not just exact words.
Vector Database
Storage optimized for finding the embeddings most similar to a query's embedding, quickly, at scale.
Retrieval
The step of actually searching the knowledge source and returning the most relevant chunks.
Reranking
An optional second pass that reorders retrieved results using a more precise (and more expensive) model.
Context
The retrieved chunks, assembled into the prompt the model will actually read.
LLM
The model that reads the question and the retrieved context, and generates the final answer.

A Real-World Example

A company has thousands of internal documents — policies, handbooks, and process guides. An employee asks an internal assistant:

"What is our parental leave policy?"

A RAG system handling this:

  1. Receives the question.
  2. Searches the internal knowledge base for relevant content.
  3. Retrieves the sections of the employee handbook that mention parental leave.
  4. Sends those retrieved sections to the LLM as context, along with the original question.
  5. Generates an answer grounded in the retrieved policy text.

Important

The model is not querying the company database by itself. Retrieval is a separate, explicit step that happens before generation — the model only ever sees the text it was handed.

Code Example

This is a minimal illustration of the pattern, not a production implementation:

rag_example.py
def answer_question(question, retriever, llm):
    # 1. Retrieve documents relevant to the question
    documents = retriever.search(question, top_k=3)

    # 2. Build context from the retrieved documents
    context = "\n\n".join(doc.text for doc in documents)

    # 3. Send the question and context to the model together
    prompt = f"""Answer the question using only the context below.

Context:
{context}

Question: {question}
Answer:"""

    return llm.generate(prompt)

Real systems add a chunking strategy, an embedding model, a vector database, metadata filtering, and often a reranking step — but the core pattern stays the same: retrieve, then generate.

Common Mistakes

  • Poor chunking

    Chunks that are too large dilute relevance; chunks that are too small lose context. Both hurt retrieval quality.

  • Retrieving irrelevant documents

    If the retriever returns weakly related content, the model will generate an answer grounded in the wrong information.

  • Retrieving too much context

    Stuffing the prompt with excessive retrieved text increases cost and can bury the passage that actually matters.

  • Ignoring metadata

    Filtering by source, date, or permissions is often as important as semantic similarity for returning the right content.

  • Assuming vector search is always enough

    Keyword and hybrid search often outperform pure vector search for exact terms, IDs, or rare vocabulary.

  • Skipping evaluation

    Without measuring retrieval quality and answer accuracy, it's hard to know whether changes actually improve the system.

  • Confusing retrieval problems with LLM problems

    A wrong answer is often caused by bad retrieval, not a weak model — the model can only work with what it was given.

  • Failing to inspect retrieved context

    The fastest way to debug a bad RAG answer is to look at what was actually retrieved, not just the final output.

Interview Question

What is RAG and why would you use it?

RAG, or Retrieval-Augmented Generation, retrieves relevant information from an external knowledge source and provides it to a language model as context before it generates a response. You'd use it when an application needs to answer using information the model wasn't trained on — private company data, frequently changing information, or domain-specific knowledge — without retraining or fine-tuning the model itself. It also reduces hallucination risk by grounding the answer in retrieved, verifiable content instead of relying purely on what the model memorized.

What an interviewer may ask next

  • Why not just fine-tune the model instead?
  • What role do embeddings play in retrieval?
  • What is chunking, and why does it matter?
  • Why would you add a reranking step?
  • How would you evaluate whether a RAG system is actually working well?

Explain It in 30 Seconds

RAG is a pattern where you retrieve relevant external information and give it to a language model as context before it generates a response. It lets an application answer using private, current, or domain-specific knowledge without requiring the model itself to memorize that information — and because the answer is grounded in retrieved content, it's less likely to be a confident-sounding guess.

On this page