RAG Application
Apply the RAG lessons to a working ingestion and retrieval architecture — from documents to a grounded answer.
What You Will Build
An illustrative RAG pipeline: an ingestion path that chunks and embeds documents into a vector store, and a query path that retrieves relevant chunks, builds context, and generates a grounded answer with citations. This project applies the RAG, Chunking, Embeddings, Retrieval, Reranking, and RAG Evaluation lessons rather than re-teaching them.
Learning Objectives
Understand why ingestion and query are separate pipelines with different requirements
Apply chunking and embedding to prepare documents for retrieval
Build a query path that retrieves, reranks, and constructs context
Understand why a RAG answer needs grounding checks and citations
Prerequisites
Concepts Used
Architecture
User Query
Starting PointOften retrieves poorly on its own.
Query Transformation
If NeededReshapes the query before retrieval.
Hybrid Retrieval
Access-FilteredFiltered by metadata so tenants stay separate.
Reranking
Narrows the SetOrders candidates before context is built.
Context Builder
Top Chunks OnlyEach tagged with its source for citation.
LLM
Answers from ContextInstructed to use only what was retrieved.
Answer + Citations
Grounding CheckedFlagged or rejected if not grounded.
Step 1 — Build the Ingestion Path
What are we doing? Turning source documents into searchable chunks with embeddings and metadata. Why? Retrieval can only find what was actually indexed — this step runs ahead of any query, and can be slow and asynchronous since no user is waiting on it.
Documents
Raw SourceWhat retrieval can eventually find.
Parsing
Extracts TextTurns source files into usable text.
Chunking
Splits It UpBreaks text into retrievable pieces.
Embedding
Vector FormMakes each chunk comparable by meaning.
Metadata Tagging
Source + TenantEnables access-filtered retrieval later.
Vector Store
Ready to QueryIndexed ahead of any user request.
def ingest(document, source_id, tenant_id):
chunks = chunk_text(document.text, max_tokens=300, overlap=50)
for chunk in chunks:
embedding = embed(chunk.text)
vector_store.upsert(
id=chunk.id,
embedding=embedding,
text=chunk.text,
metadata={"source": source_id, "tenant": tenant_id},
)Step 2 — Build the Query Path
What are we doing? Turning a user's question into a set of relevant, access-filtered chunks. Why? A raw query often retrieves poorly on its own, and unfiltered retrieval can leak content across tenants. How it works: transform the query if needed, retrieve with hybrid search filtered by metadata, then rerank the results.
def retrieve(query, tenant_id, top_k=5):
candidates = hybrid_search(query, filter={"tenant": tenant_id}, limit=20)
return rerank(query, candidates)[:top_k]Step 3 — Build the Context
What are we doing? Assembling the retrieved chunks into the prompt the model actually sees. Why? Too much context wastes tokens and can bury the relevant passage; too little starves the model of what it needs. How it works: include only the top reranked chunks, each tagged with its source for later citation.
Tip
Tag each chunk in the context with an identifier the model can reference — this is what makes citations possible on the way out.
Step 4 — Generate and Cite
What are we doing? Generating the answer and validating that it's actually grounded in the retrieved context. Why? A model can still produce an ungrounded, hallucinated answer even with good context available. How it works: instruct the model to answer only from the provided context and cite which chunk supports each claim; flag or reject answers that don't reference any retrieved content.
Generate answer
Return directly
Risk: confident but ungrounded
Generate answer
Verify claims against context
Flag or reject if ungrounded
Step 5 — Handle Failure and Add Observability
What are we doing? Deciding what happens when retrieval fails or returns nothing useful, and what to track. Why? A wrong answer in RAG is often a retrieval failure, not a generation failure — you need visibility to tell the difference.
Skipping metadata filtering for access control
Permission enforcement has to happen at retrieval time, not by hoping the model omits content it shouldn't use.
No fallback when retrieval returns nothing relevant
The system should say it doesn't know rather than forcing an answer with no useful context.
Treating ingestion and query as one pipeline
They have different performance requirements and failure modes.
Never running RAG evaluation after launch
A system that looked good in a demo can regress silently as content or settings change.
Assuming citations guarantee correctness
A citation shows what the model referenced, not that its interpretation of that source was accurate.
Challenges
Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.
Challenge 1: Improve retrieval quality
Experiment with chunk size and overlap and observe the effect on retrieved relevance.
Challenge 2: Add metadata filtering
Add tenant- or permission-based filtering so retrieval respects access boundaries.
Challenge 3: Add reranking
Add a reranking step after initial retrieval and compare answer quality with and without it.
Challenge 4: Measure retrieval quality
Build a small golden test set and measure retrieval relevance and groundedness, as covered in RAG Evaluation.
Design Review
Before moving on, think through these questions the way a reviewer would.
What would you change if this needed to serve 10x the document volume?
Where is the latency bottleneck in the query path?
What happens if the vector database is temporarily unavailable?
How would you prevent one tenant's documents from appearing in another tenant's answers?
How would you know if a change to chunking made retrieval worse?
Interview Questions
How would you design a production RAG system for an enterprise knowledge base?
I'd split it into two pipelines. Ingestion runs ahead of time: parsing, chunking, embedding, and tagging metadata like source and permissions, written to a vector store — this can be async and batched. The query pipeline runs per request: transform the query if needed, retrieve with hybrid search filtered by access metadata, rerank, then build context for the model. On the way out, I'd validate the answer is actually grounded in what was retrieved and require citations, and I'd run RAG evaluation continuously, since a wrong answer is often bad retrieval, not a bad model.
- Separate ingestion and query pipelines
- Access control enforced at retrieval time
- Groundedness validation and citations
- Continuous evaluation
A user reports a wrong answer. How do you figure out whether it was a retrieval problem or a generation problem?
I'd inspect what was actually retrieved for that query. If the relevant content wasn't in the retrieved set at all, that's a retrieval failure — worth checking chunking, embedding quality, or filtering. If the relevant content was retrieved but the model still answered incorrectly or ignored it, that's a generation or grounding failure, worth checking the context builder and the model's instructions.
- Separate evaluation of retrieval vs. generation
- Inspect the actual retrieved chunks
- Distinguishes the fix needed
How would you prevent tenant data from leaking across a shared vector store?
I'd tag every chunk with tenant metadata at ingestion time and filter every retrieval query by the requesting tenant — enforced in the retrieval code itself, not left to the model to figure out. I wouldn't rely on the model to withhold content it was already given.
- Metadata filtering enforced at retrieval time
- Never rely on model behavior for access control
Explain It in 30 Seconds
This project applies the RAG lessons to a working pipeline: ingestion parses, chunks, embeds, and tags documents ahead of time, while the query path transforms the question, retrieves with hybrid search filtered by access metadata, reranks, and builds context for the model. Generation includes a grounding check and citations, since a model can still hallucinate even with good context. A wrong answer is often a retrieval failure, not a generation failure — which is why continuous evaluation matters.