Observability
Observability is the ability to see what an AI system is doing in production through logs, traces, and metrics.
Prerequisites
Traditional vs. AI Observability
Traditional observability tells you whether a system is up, how fast it's responding, and where errors are occurring — standard metrics, logs, and traces. AI systems need all of that, plus a second layer: is the system actually producing good outputs? A RAG pipeline can be fast, reliable, and completely healthy by traditional metrics while consistently giving wrong answers because retrieval quality quietly degraded.
Application
Entry PointWhere a request first enters the system.
Gateway
Routing LayerCan be healthy while quality quietly degrades.
Retrieval
Quality RiskFast and reliable, but not necessarily correct.
Tools
External CallsAnother stage that can silently underperform.
LLM
Output QualityStandard metrics say nothing about correctness.
Metrics / Logs / Traces at every stage
Second LayerStandard health metrics plus output quality.
Key Idea
AI observability requires both system observability (is it running correctly?) and quality observability (is it producing good results?) — traditional tooling only covers the first.
What to Track
- Metrics
- Latency, throughput, error rate, token usage, cost, retrieval quality signals, and tool failure rates — aggregated numbers you can alert on and trend over time.
- Logs
- Per-request detail: which model handled it, what was retrieved, which tool was called, what errors occurred — the detail you need to investigate a specific incident.
- Traces
- The full path a single request took across application, gateway, retrieval, tools, and the model — essential for understanding where time and cost went in a multi-step request.
- Quality signals
- Groundedness, relevance, correctness, and user feedback — the metrics that answer whether the system is actually helpful, not just whether it's running.
A Practical Example
A RAG system's average latency and error rate look completely normal, but user satisfaction quietly drops. Traditional observability alone wouldn't catch this — it takes tracking retrieval quality and groundedness specifically to notice that a recent change to chunking made retrieval less precise, even though nothing "broke" in the traditional sense.
Warning
Be deliberate about what gets logged. Full prompts and responses often contain sensitive user content, and logging everything by default is a common way that content leaks into logs unnecessarily.
Common Mistakes
Only tracking traditional system metrics
Latency and error rate can look perfectly healthy while output quality quietly degrades — quality signals need their own tracking.
No tracing across a multi-step request
Without a trace spanning retrieval, tools, and the model, it's hard to tell which stage of a slow or failed request was actually the bottleneck.
Logging full prompts and responses indiscriminately
This is a common way sensitive content ends up retained in logs longer, and more broadly, than intended.
Not tracking token usage and cost per request
Without this, cost problems are only discovered when the bill arrives, not when the underlying usage pattern actually changes.
Treating observability as something added after launch
Retrofitting tracing and metrics into an already-complex multi-stage system is much harder than building them in as the system is developed.
Interview Question
How would you design observability for a production RAG or agent system?
I'd track two layers, not just one: traditional system observability — latency, throughput, error rate, token usage, cost — and AI-specific quality observability — retrieval relevance, groundedness, tool failure rates, and user feedback. Traditional metrics alone can look completely healthy while the system quietly produces worse answers, for example if a chunking change degrades retrieval precision without causing any errors. I'd use tracing across the full request path — application, gateway, retrieval, tools, model — to see where time and cost actually go in a multi-step request, and I'd be deliberate about what gets logged, since full prompts and responses often contain sensitive content that shouldn't be captured indiscriminately.
What an interviewer may ask next
- Why can a RAG system look healthy on traditional metrics while actually producing worse answers?
- What does tracing add that logs and metrics alone don't?
- What are the risks of logging full prompts and responses?
Explain It in 30 Seconds
AI observability needs two layers: traditional system observability — latency, error rate, throughput, cost — and quality observability — groundedness, retrieval relevance, tool failures, user feedback — since a system can look perfectly healthy by traditional metrics while quietly producing worse answers. Tracing across the full request path, from gateway through retrieval, tools, and the model, is what lets you see where time and cost actually go in a multi-step request. Logging needs care too, since full prompts and responses often contain sensitive content.