Context Window
The context window is the maximum amount of text a model can consider at once when generating a response.
Prerequisites
What Is a Context Window?
A model can only "see" a limited number of tokens at once — the system prompt, the conversation so far, any retrieved documents, and the response it's generating, all combined. That limit is the context window, measured in tokens.
System Prompt
Standing RulesThe application's instructions to the model.
Conversation History
Prior TurnsAs much of the past conversation as fits.
Retrieved Context
Added DocumentsRetrieved content or tool results, if any.
User Message
Current TurnWhat the user just asked.
Model Output
Also CountedThe response itself also uses up context space.
Key Idea
The context window is a shared budget. Everything the model reads and everything it generates draws from the same total token limit.
Why It Matters
- Long conversations eventually exceed the window — older messages have to be dropped, summarized, or truncated to make room for new ones.
- RAG systems compete for the same budget — retrieved chunks, the system prompt, and the conversation all share one window, so retrieving too much content can crowd out other important context.
- A bigger context window is not automatically better — the model still has to find the relevant part of everything you give it, and stuffing in unnecessary text can dilute what actually matters.
- Cost and latency scale with how much context is sent, not just the output — providers typically charge for both input and output tokens.
A Real-World Example
In a long chat session, a user might reference something they said many messages ago. If that message has aged out of the context window — because newer messages pushed it out to stay within the limit — the model genuinely no longer has access to it, and will answer as if it was never said. This is why long-running assistants often summarize or selectively carry forward only the most relevant earlier context, rather than keeping the entire raw history.
Common Mistakes
Assuming a bigger context window always improves answers
A larger window means the model can accept more input, not that it will find or use the relevant parts of it equally well.
Forgetting that output tokens count against the same limit
The context window covers everything read and everything generated combined, not just the input.
Sending the full conversation history indefinitely
Without truncation or summarization, a long-running conversation will eventually exceed the window and start silently losing earlier content.
Retrieving as much context as possible "to be safe"
Over-retrieving in a RAG system can crowd out the system prompt and conversation, and can bury the passage that actually answers the question.
Interview Question
What is a context window, and why does it matter when building with LLMs?
The context window is the maximum number of tokens a model can consider at once, covering the system prompt, conversation history, any retrieved context, and the response it generates — all sharing one budget. It matters because everything you send and everything the model produces draws from that same limit: long conversations eventually push out earlier messages, and RAG systems have to balance how much retrieved content to include against the system prompt and conversation. A bigger window isn't automatically better either — the model still has to locate what's relevant within everything it's given.
What an interviewer may ask next
- What happens to earlier messages in a long conversation once the context window fills up?
- Why might retrieving too much content in a RAG system actually hurt answer quality?
- Does output count against the context window, or only input?
Explain It in 30 Seconds
The context window is the total number of tokens a model can consider at once — system prompt, conversation history, retrieved context, and its own output all draw from that same shared budget. Once you exceed it, older content has to be dropped, summarized, or truncated. A larger window lets you send more, but doesn't guarantee the model uses it well — irrelevant or excessive context can still bury what actually matters.