RAG Architecture
RAG architecture describes the end-to-end system design for retrieval, ranking, and generation in a production application.
Prerequisites
From Pattern to System
The RAG lesson explains the core pattern: retrieve relevant content, then generate an answer using it. This lesson assumes that pattern and asks a different question — what does it take to run that pattern as a reliable, production system, made up of an ingestion pipeline that prepares data in advance and a query pipeline that runs on every request?
Key Idea
RAG architecture has two independent pipelines with very different failure modes and performance requirements: ingestion (runs occasionally, can be slow) and query (runs on every request, has to be fast).
Minimal RAG vs. Production RAG
App
Retriever
LLM
Response
API + Auth
Query Transformation
Hybrid Retrieval
Metadata Filtering
Reranking
Context Builder
LLM (via Gateway)
Validation / Citations
Observability
Each added stage in the production version exists to solve a problem the minimal version doesn't handle: query transformation improves recall for poorly-phrased queries, metadata filtering enforces access control, reranking improves precision, and validation catches ungrounded answers before they reach the user. None of these are required to build a working RAG demo — they become necessary as a system moves toward production.
The Ingestion Side
Source Documents
Raw ContentDocuments before any processing.
Parsing
Extracts TextTurns raw files into usable text.
Chunking
Splits ContentSmaller pieces sized for retrieval.
Embedding
Vector FormEach chunk converted for similarity search.
Metadata Tagging
Enables FilteringAlso enforces access control at query time.
Index / Vector Database
Query-ReadyWhat retrieval reads from at query time.
- Data freshness — documents change; the pipeline needs a strategy for re-ingesting updated or deleted content, not just adding new documents.
- Access-aware indexing — permission metadata has to be captured at ingestion time, since it can't be reconstructed at query time.
- Idempotency — re-running ingestion on the same document shouldn't silently duplicate it in the index.
- This pipeline can run asynchronously, in batches, and can tolerate being slower than the query path — it doesn't have a user waiting on it in real time.
Failure Modes
- Vector database or index unavailable — the query pipeline has no candidates to retrieve; the system needs a defined degraded response, not a hard crash.
- Retrieval returns irrelevant results — often the actual cause of a wrong answer, not the model itself.
- Reranking or embedding model latency — an added pipeline stage is an added point of slowness on the critical path.
- Stale index — ingestion falling behind means the system answers confidently from outdated information.
- Ungrounded generation — the model answers beyond what was actually retrieved, which is exactly what RAG evaluation and citation checks are meant to catch.
Warning
A wrong answer from a RAG system is often a retrieval problem, not a model problem — architecture and evaluation should make it possible to tell which stage actually failed.
Common Mistakes
Treating ingestion and query as one pipeline
They have different performance requirements and failure modes — ingestion can be slow and batched, query has to be fast and reliable per request.
No fallback when retrieval returns nothing relevant
The system should have a defined behavior — like saying it doesn't know — rather than forcing the model to generate an answer with no useful context.
Skipping metadata filtering for access control
Permission enforcement has to happen at retrieval time, not by hoping the model omits content it shouldn't use.
No re-ingestion strategy for updated documents
Without one, the index silently drifts out of date relative to the source of truth.
Never running RAG evaluation in production
A system that looked good in a demo can regress silently as content, prompts, or retrieval settings change.
Putting the entire pipeline inline in the request path
Query transformation, retrieval, reranking, and generation each add latency — the architecture should identify what can be parallelized or cached.
Interview Question
How would you design a production RAG system for an enterprise knowledge base?
I'd split it into two pipelines with different requirements. Ingestion runs ahead of time: parsing documents, chunking them, generating embeddings, tagging metadata like source and permissions, and writing to an index — this can be asynchronous and batched, and needs a strategy for re-ingesting updated or deleted content. The query pipeline runs per request and needs to be fast: query transformation to handle poorly-phrased questions, hybrid retrieval filtered by access-control metadata, reranking to improve precision, then a context builder that assembles what the model actually sees. On the way out, I'd validate that the answer is actually grounded in what was retrieved before returning it, and I'd run RAG evaluation continuously, because a wrong answer is often bad retrieval, not a bad model, and evaluation is how you tell the difference.
What an interviewer may ask next
- Why should ingestion and query be treated as separate pipelines with different requirements?
- How would you enforce access control so users only retrieve documents they're allowed to see?
- What would you do if the vector database or index became unavailable?
- How would you diagnose whether a wrong answer was caused by retrieval or generation?
Explain It in 30 Seconds
RAG architecture splits into two pipelines: ingestion, which parses, chunks, embeds, and indexes documents ahead of time and can run asynchronously, and query, which runs per request and needs to be fast — query transformation, hybrid retrieval with metadata filtering for access control, reranking, and a context builder feeding the model. Production RAG adds these stages to solve specific problems minimal RAG doesn't handle: recall, access control, precision, and grounding. A wrong answer is often a retrieval failure, not a model failure, which is why continuous evaluation matters.