AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Advanced7 min read

RAG Architecture

RAG architecture describes the end-to-end system design for retrieval, ranking, and generation in a production application.

Prerequisites

From Pattern to System

The RAG lesson explains the core pattern: retrieve relevant content, then generate an answer using it. This lesson assumes that pattern and asks a different question — what does it take to run that pattern as a reliable, production system, made up of an ingestion pipeline that prepares data in advance and a query pipeline that runs on every request?

Key Idea

RAG architecture has two independent pipelines with very different failure modes and performance requirements: ingestion (runs occasionally, can be slow) and query (runs on every request, has to be fast).

Minimal RAG vs. Production RAG

Minimal RAG

App

Retriever

LLM

Response

Production RAG (query time)

API + Auth

Query Transformation

Hybrid Retrieval

Metadata Filtering

Reranking

Context Builder

LLM (via Gateway)

Validation / Citations

Observability

Each added stage in the production version exists to solve a problem the minimal version doesn't handle: query transformation improves recall for poorly-phrased queries, metadata filtering enforces access control, reranking improves precision, and validation catches ungrounded answers before they reach the user. None of these are required to build a working RAG demo — they become necessary as a system moves toward production.

The Ingestion Side

Source Documents

Raw Content

Documents before any processing.

processed by

Parsing

Extracts Text

Turns raw files into usable text.

split via

Chunking

Splits Content

Smaller pieces sized for retrieval.

converted via

Embedding

Vector Form

Each chunk converted for similarity search.

tagged with

Metadata Tagging

Enables Filtering

Also enforces access control at query time.

stored in

Index / Vector Database

Query-Ready

What retrieval reads from at query time.

Ingestion pipeline (runs ahead of time)
  • Data freshness — documents change; the pipeline needs a strategy for re-ingesting updated or deleted content, not just adding new documents.
  • Access-aware indexing — permission metadata has to be captured at ingestion time, since it can't be reconstructed at query time.
  • Idempotency — re-running ingestion on the same document shouldn't silently duplicate it in the index.
  • This pipeline can run asynchronously, in batches, and can tolerate being slower than the query path — it doesn't have a user waiting on it in real time.

Failure Modes

  • Vector database or index unavailable — the query pipeline has no candidates to retrieve; the system needs a defined degraded response, not a hard crash.
  • Retrieval returns irrelevant results — often the actual cause of a wrong answer, not the model itself.
  • Reranking or embedding model latency — an added pipeline stage is an added point of slowness on the critical path.
  • Stale index — ingestion falling behind means the system answers confidently from outdated information.
  • Ungrounded generation — the model answers beyond what was actually retrieved, which is exactly what RAG evaluation and citation checks are meant to catch.

Warning

A wrong answer from a RAG system is often a retrieval problem, not a model problem — architecture and evaluation should make it possible to tell which stage actually failed.

Common Mistakes

  • Treating ingestion and query as one pipeline

    They have different performance requirements and failure modes — ingestion can be slow and batched, query has to be fast and reliable per request.

  • No fallback when retrieval returns nothing relevant

    The system should have a defined behavior — like saying it doesn't know — rather than forcing the model to generate an answer with no useful context.

  • Skipping metadata filtering for access control

    Permission enforcement has to happen at retrieval time, not by hoping the model omits content it shouldn't use.

  • No re-ingestion strategy for updated documents

    Without one, the index silently drifts out of date relative to the source of truth.

  • Never running RAG evaluation in production

    A system that looked good in a demo can regress silently as content, prompts, or retrieval settings change.

  • Putting the entire pipeline inline in the request path

    Query transformation, retrieval, reranking, and generation each add latency — the architecture should identify what can be parallelized or cached.

Interview Question

How would you design a production RAG system for an enterprise knowledge base?

I'd split it into two pipelines with different requirements. Ingestion runs ahead of time: parsing documents, chunking them, generating embeddings, tagging metadata like source and permissions, and writing to an index — this can be asynchronous and batched, and needs a strategy for re-ingesting updated or deleted content. The query pipeline runs per request and needs to be fast: query transformation to handle poorly-phrased questions, hybrid retrieval filtered by access-control metadata, reranking to improve precision, then a context builder that assembles what the model actually sees. On the way out, I'd validate that the answer is actually grounded in what was retrieved before returning it, and I'd run RAG evaluation continuously, because a wrong answer is often bad retrieval, not a bad model, and evaluation is how you tell the difference.

What an interviewer may ask next

  • Why should ingestion and query be treated as separate pipelines with different requirements?
  • How would you enforce access control so users only retrieve documents they're allowed to see?
  • What would you do if the vector database or index became unavailable?
  • How would you diagnose whether a wrong answer was caused by retrieval or generation?

Explain It in 30 Seconds

RAG architecture splits into two pipelines: ingestion, which parses, chunks, embeds, and indexes documents ahead of time and can run asynchronously, and query, which runs per request and needs to be fast — query transformation, hybrid retrieval with metadata filtering for access control, reranking, and a context builder feeding the model. Production RAG adds these stages to solve specific problems minimal RAG doesn't handle: recall, access control, precision, and grounding. A wrong answer is often a retrieval failure, not a model failure, which is why continuous evaluation matters.

On this page