Model Evaluation in Production
Production model evaluation tracks quality continuously against real traffic, not just a one-time benchmark before launch.
Prerequisites
Overview
A model that passed evaluation before launch can still degrade in production — a provider silently updates a model version, usage patterns shift, or edge cases the original test set never covered start showing up.
Where It Fits
Live Traffic Sample
Automated Scoring
Human Spot-Check
Quality Alert
Key Points
- Continuous sampling
- A sample of real production requests is scored on an ongoing basis, not just a static test set run once.
- Drift detection
- A drop in quality metrics over time can signal a silent provider model update or a shift in the kinds of requests coming in.
- Human-in-the-loop review
- Automated scoring is often paired with periodic human review of a sample, since some quality issues are hard to detect programmatically.
Interview Question
A model that passed evaluation at launch is producing worse answers three months later. What could explain that?
A few possibilities: the provider silently updated the underlying model version, the mix of real user requests has shifted away from what the original test set covered, or upstream data (like retrieved documents) has drifted. This is exactly why production evaluation needs to run continuously against live traffic, not just once before launch.
Explain It in 30 Seconds
Production model evaluation continuously scores a sample of live traffic rather than relying on a one-time pre-launch benchmark, since silent provider updates or shifting usage patterns can degrade quality after launch.
Real-World Stack
Technologies commonly used to implement this in production.