Prompt Evaluation
Evaluating a prompt means testing it against a fixed set of cases before and after any change, not just eyeballing a few outputs.
Prerequisites
Overview
It’s easy to tweak a prompt, try it a couple of times, and assume it’s better. Prompt evaluation replaces that with a fixed test set run automatically, so a change’s actual effect — including regressions on cases that used to work — is visible before it ships.
Where It Fits
Fixed Test Cases
Current & Candidate Prompts
Both run against the same test cases.
Compare Results
Key Points
- Fixed test set
- A stable set of representative inputs and expected qualities, reused for every prompt change rather than invented ad hoc.
- Regression detection
- Evaluation catches cases that used to work correctly but broke because of a new change — easy to miss by spot-checking.
- Qualitative and quantitative checks
- Some criteria (format correctness) are checkable programmatically; others (tone, helpfulness) may need a scoring rubric or model-based grading.
Interview Question
Why isn’t trying a new prompt a few times and reading the outputs a sufficient way to validate it?
A handful of manual tries can’t reveal regressions on the range of inputs the prompt needs to handle in production, and human judgment on a small sample is inconsistent. A fixed evaluation set run automatically catches both new improvements and unintended regressions across a representative range of cases, not just the ones a person happened to try.
Explain It in 30 Seconds
Prompt evaluation runs a fixed set of test cases against a candidate prompt automatically, catching regressions and measuring real improvement — replacing the unreliable habit of judging a prompt change from a few manual tries.
Real-World Stack
Technologies commonly used to implement this in production.