LLM Observability Tools
Tools like Langfuse, LangSmith, and Arize Phoenix trace every model and tool call so a developer can see exactly what happened.
Prerequisites
Overview
Generic application logs weren’t built for LLM-specific concerns — token usage, prompt versions, retrieved context, tool call arguments. Dedicated LLM observability tools capture all of that as structured traces purpose-built for debugging AI systems.
Where It Fits
Request
Prompt Version Used
Model Call + Tokens
Trace Stored
Key Points
- Structured for LLM concerns
- These tools natively understand prompts, token counts, and tool calls, unlike generic application logging.
- Cost per request
- Token-level tracking makes it possible to see cost per user, per feature, or per prompt version.
- Prompt/version correlation
- A trace usually records which prompt version and model produced a given response, essential for debugging a regression.
Interview Question
Why do teams use a dedicated LLM observability tool instead of their existing application logging?
General application logs weren’t built to capture prompt versions, token usage, retrieved context, or tool call arguments in a structured, queryable way. A dedicated tool like Langfuse or LangSmith treats these as first-class fields, making it possible to trace a specific bad response back to the exact prompt version and inputs that produced it.
Explain It in 30 Seconds
LLM observability tools like Langfuse, LangSmith, and Arize Phoenix capture prompt versions, token usage, and tool calls as structured traces purpose-built for debugging AI systems, which generic application logging doesn’t naturally provide.
Real-World Stack
Technologies commonly used to implement this in production.