AI Memory
Memory lets an agent or application retain information across steps or sessions, instead of treating every interaction as if nothing came before it.
Prerequisites
Context Window vs. Memory
It's tempting to think a bigger context window solves memory — just put everything in the prompt. But the context window is the amount of text a model can consider for one request. Memory is something else: the ability to retain and reuse information across separate requests, conversations, or sessions.
Important
Putting more text into the context window is not automatically "memory." Real memory means deciding what to store, how to retrieve it later, and making it available in a future interaction the model wouldn't otherwise have any information about.
Current Context
This Request OnlyEverything the model can see for one call.
+ Stored Information
Retained Across TimeWhat was decided worth saving from earlier.
Useful Future Context
Real MemoryAvailable to a session that never saw it originally.
Kinds of Memory
- Short-term / conversational memory
- The recent back-and-forth in an ongoing conversation, usually just carried directly in the context window.
- Long-term memory
- Information saved outside a single conversation — often in a vector database — and retrieved when relevant in a future session.
- Semantic memory
- General facts or preferences worth retaining, like a user’s stated preferences or established facts about a project.
- Episodic memory
- A record of specific past events or interactions — what happened last time, not just a general fact.
How It Actually Works
In practice, "giving a system memory" usually means: deciding what is worth storing (often via summarization, so you keep the gist rather than every raw message), storing it somewhere retrievable, and retrieving the relevant pieces — frequently with the same retrieval pattern used in RAG — when they become useful again. Common uses include remembering user preferences and tracking the state of a longer, ongoing task.
Common Mistakes
Confusing a long context window with memory
A large context window helps within one request; it does not, by itself, make information available in a future session.
Storing everything without summarization
Saving every raw message indefinitely gets expensive and noisy — most systems benefit from summarizing what actually matters.
No retrieval strategy for stored memory
Storing information is only half the problem — without a way to find the relevant piece later, stored memory just accumulates unused.
Interview Question
What is the difference between a context window and memory?
The context window is the amount of text a model can consider for a single request — it resets with every new, independent call. Memory is the ability to retain and reuse information across separate requests, conversations, or sessions, which requires actually deciding what to store, where to store it, and how to retrieve the relevant pieces later. Just putting more text into a prompt isn't memory — real memory involves storage and retrieval that persists beyond a single interaction.
What an interviewer may ask next
- What is the difference between short-term and long-term memory in this context?
- Why would you summarize stored memory instead of keeping raw history?
- How does memory retrieval relate to retrieval in a RAG system?
Explain It in 30 Seconds
Memory lets a system retain and reuse information across separate conversations or sessions, which is different from the context window — the amount of text available to a single request. Real memory means deciding what to store, often by summarizing, and having a way to retrieve the relevant piece later, similar to retrieval in a RAG system. Putting more text into one prompt is not the same thing as giving a system memory.