Reflection
Reflection is an agent reviewing its own output or actions to catch errors and improve before continuing.
Prerequisites
What Is Reflection?
By default, an agent moves forward after each step in its loop without necessarily double-checking its own work. Reflection adds an explicit step where the agent — or a separate evaluation call — reviews the most recent output or action, checking whether it actually made progress, before deciding whether to continue, retry, or change approach.
Action
Not Yet CheckedWhat the agent just did in the loop.
Result
Raw OutcomeWhat happened, before anyone reviews it.
Reflection: Did This Work?
Explicit ReviewThe agent, or a separate call, checks progress.
Continue / Retry / Adjust Plan
Informed DecisionBased on the review, not just moving forward blindly.
Key Idea
Reflection is what lets an agent catch its own mistakes mid-task, instead of only discovering a failure at the very end.
How Reflection Is Implemented
- Self-critique — asking the model to evaluate its own previous output against the goal, in a separate step, before proceeding.
- A dedicated evaluator — a separate check (sometimes a smaller, cheaper model, sometimes rule-based) that judges progress independently of the agent's own self-assessment.
- Retry with feedback — when reflection identifies a problem, the agent gets another attempt, informed by what specifically went wrong the first time.
Warning
A model reflecting on its own output shares the same blind spots that produced the output in the first place — reflection reduces certain kinds of errors, but it's not a guarantee of correctness.
Common Mistakes
Assuming reflection catches every kind of error
Self-critique shares the model's own blind spots — it helps with some errors, but isn't a substitute for independent evaluation or guardrails.
Adding reflection to every single step regardless of need
Reflection adds latency and cost on every pass — it's most valuable at meaningful checkpoints, not necessarily after every minor action.
No limit on retry attempts after a failed reflection check
Without a cap, an agent stuck in a reflect-retry cycle can loop as unboundedly as one with no reflection at all.
Treating reflection as a replacement for the evaluator role in the agent architecture
Reflection is one technique an evaluator component might use — it's not automatically the same thing as having genuine independent evaluation.
Interview Question
What is reflection in an agent system, and what are its limits?
Reflection is an explicit step where an agent — or a separate check — reviews its own recent output or action against the goal before deciding whether to continue, retry, or adjust its approach, instead of just moving forward blindly after every step. It's usually implemented as self-critique, a dedicated evaluator judging progress independently, or a retry informed by specific feedback about what went wrong. The main limit is that a model reflecting on its own output shares the same blind spots that produced the output in the first place — it reduces certain kinds of errors, but it's not a guarantee of correctness, and it still needs a bounded retry limit so a reflect-retry cycle doesn't loop indefinitely.
What an interviewer may ask next
- Why doesn't reflection guarantee an agent's output is correct?
- How does self-critique differ from using a dedicated, independent evaluator?
- Why does reflection still need a retry limit?
Explain It in 30 Seconds
Reflection is an explicit step where an agent reviews its own recent output or action against the goal before continuing, retrying, or adjusting its approach — catching mistakes mid-task instead of only at the end. It's implemented as self-critique, a separate independent evaluator, or a retry informed by specific feedback. Its main limit is that self-critique shares the same blind spots as the output it's reviewing, so it reduces certain errors without guaranteeing correctness, and still needs a bounded retry limit.