Document Ingestion for RAG
Ingestion keeps a RAG system in sync with its source documents — detecting new, updated, and deleted content over time.
Prerequisites
Overview
A RAG system’s knowledge is only as fresh as its ingestion process. Ingestion isn’t a one-time load — it’s an ongoing job that detects new documents, re-processes changed ones, and removes deleted ones from the index.
Where It Fits
Source System
Change Detection
Re-process Changed Docs
Vector Index Updated
Key Points
- Incremental sync
- Re-processing only changed documents, rather than the entire corpus, keeps ingestion fast and cheap as content grows.
- Deletion handling
- A deleted source document needs its chunks removed from the index too, or retrieval can surface stale content.
- Scheduling
- Ingestion typically runs on a schedule or is triggered by an event (a webhook, a Kafka message) rather than manually.
Interview Question
What happens to a RAG system’s answers if ingestion silently stops running?
The index gradually goes stale — new documents never appear, and updates to existing ones aren’t reflected — while the system keeps confidently answering from outdated content, since nothing in generation itself signals that retrieval is working from old data. Ingestion needs its own monitoring, separate from generation quality checks.
Explain It in 30 Seconds
Document ingestion is an ongoing process, not a one-time load — it detects new, changed, and deleted source documents and keeps a RAG system’s index in sync, and needs its own monitoring since generation quality alone won’t reveal a stale index.
Real-World Stack
Technologies commonly used to implement this in production.