Chunking
Chunking splits documents into smaller pieces so they can be embedded and retrieved effectively, instead of retrieving an entire document at once.
Prerequisites
Why Documents Are Chunked
A language model has a limited context window, and retrieval works better when it can return a focused, relevant piece of text rather than an entire document. Chunking is the step that splits documents into smaller pieces before they are embedded and stored, so retrieval can return just the part that actually answers a question.
Document
Raw InputThe original file before it's split into pieces.
Chunks
Split Into PiecesSmaller pieces sized for focused retrieval, not a whole document.
Embeddings
Vector FormEach chunk converted into a vector for similarity search.
Retrieval
Matched ResultsFinds the chunks that best answer the question.
The Size Tradeoff
Missing surrounding context
Fragmented meaning
Noisy, unfocused retrieval
Wastes context window
There is no single correct chunk size — it depends on the content and the questions users actually ask. A support article and a legal contract have very different natural units of meaning, and the right chunk size for one may be wrong for the other.
Chunking Strategies
- Fixed-size chunking — split every N tokens or characters; simple, but can cut sentences or ideas in half.
- Overlap — include a small amount of shared text between consecutive chunks so context near a boundary isn’t lost.
- Semantic chunking — split at natural meaning boundaries (topic changes) rather than a fixed size.
- Structure-aware chunking — split along the document’s own structure, like headings, paragraphs, or sections.
Metadata — like the section heading, source document, or page number — is usually attached to each chunk so it can be filtered on and cited back to the user later.
Common Mistakes
Using one fixed chunk size for every document type
A support FAQ and a technical manual have different natural structures — the same chunk size rarely serves both well.
No overlap between chunks
Important context sitting right at a chunk boundary can get cut off entirely without a small amount of shared overlap.
Ignoring document structure
Splitting purely by character count can sever a heading from the content it introduces, hurting both retrieval and readability.
Interview Question
How would you choose chunk size for a RAG system?
There isn't a universal number — it depends on the content and the kinds of questions users ask. I'd start by looking at the natural structure of the documents: short, self-contained sections suggest smaller chunks; dense, interdependent content might need larger ones with overlap so context near a boundary isn't lost. I'd also lean on structure-aware or semantic chunking rather than a fixed size where the source content has clear sections, and I'd evaluate actual retrieval quality on real questions rather than picking a number up front and assuming it's right.
What an interviewer may ask next
- What goes wrong if chunks are too small? Too large?
- Why would you add overlap between chunks?
- How would you evaluate whether your chunking strategy is working?
Explain It in 30 Seconds
Chunking splits documents into smaller pieces before they’re embedded, so retrieval can return a focused, relevant piece of text instead of an entire document. Chunks that are too small lose surrounding context; chunks that are too large make retrieval noisy and waste the context window. There’s no universal chunk size — it depends on the content’s natural structure, which is why semantic or structure-aware chunking often beats a fixed size.