Caching
Caching stores previous results so repeated or similar requests can be served faster and more cheaply.
Why Caching Matters More for AI Applications
Every uncached LLM call costs real latency and, with most providers, real money. Unlike a typical API call, the expensive part isn't a database lookup — it's the model generating output token by token. Caching avoids repeating that expensive work when it isn't necessary.
Request
Incoming CallA new request that may already have an answer.
Cache Lookup
Check FirstChecked before doing any expensive work.
Return Cached Response
Fast PathSkips the model call entirely.
Call LLM
Expensive PathThe costly, token-by-token generation step.
Store Response
Cache UpdateSaved so the next matching request is a hit.
Return Response
Fresh PathFreshly generated, not from cache.
What Can Be Cached
- Response caching
- Storing the full output for an exact or near-exact repeated request — the most straightforward case.
- Embedding caching
- Storing the embedding for content that's embedded repeatedly, avoiding redundant embedding calls for identical text.
- Retrieval caching
- Caching the retrieved results for a given query, so an identical or very similar query doesn't re-run the full retrieval pipeline.
- Semantic caching
- Instead of requiring an exact match, checking whether a new query is semantically similar enough to a previously cached one to reuse its result — a more aggressive form of caching with a real risk of returning a subtly wrong answer.
- Prompt / context caching
- Some providers let you cache a large, unchanging portion of a prompt — like a long system prompt or reference document — so repeated requests don't reprocess it from scratch.
Freshness and Correctness
Caching trades some freshness for speed and cost — a cached response can become stale if the underlying data or context changes. Invalidation matters as much as caching itself: a cache with no invalidation strategy quietly serves outdated answers indefinitely.
Warning
Semantic caching is the riskiest form — two questions can be similar in wording but expect meaningfully different answers, so a cache hit doesn't guarantee correctness the way an exact match does.
Common Mistakes
Caching without an invalidation strategy
A cache with no expiration or invalidation plan will keep serving outdated answers as the underlying data changes.
Using semantic caching without understanding its risk
A cache hit based on similarity, not an exact match, can return a plausible-but-wrong answer for a subtly different question.
Caching personalized or tenant-specific responses globally
A cache key needs to include enough context — like user or tenant identity — or one user's cached response can leak to another.
Not measuring the actual cache hit rate
Without visibility into how often the cache is actually helping, it's hard to know whether the added complexity is paying for itself.
Interview Question
What can you cache in an LLM application, and what's the risk with semantic caching specifically?
You can cache full responses for repeated requests, embeddings for repeatedly embedded content, retrieval results for repeated queries, and in some providers, large unchanging portions of a prompt like a long system prompt or reference document. Semantic caching goes further, treating a new query as a cache hit if it's similar enough to a previous one, not just identical — that's the riskiest form, because two questions can be worded similarly but expect meaningfully different answers, so a cache hit doesn't guarantee correctness the way an exact match does. Whatever you cache, you need an invalidation strategy, or the cache will keep serving outdated answers as the underlying data changes.
What an interviewer may ask next
- Why is semantic caching riskier than exact-match caching?
- What happens if a cache has no invalidation strategy?
- How would you avoid one user's cached response leaking to another user?
Explain It in 30 Seconds
Caching stores previous results so repeated or similar requests skip the expensive work of generating a fresh response — you can cache full responses, embeddings, retrieval results, or large unchanging portions of a prompt. Semantic caching, which matches on similarity rather than an exact request, is the riskiest form, since a similar-sounding question can expect a meaningfully different answer. Any caching strategy needs invalidation, or it will quietly keep serving outdated answers.