Human-in-the-Loop
Human-in-the-loop keeps a person able to review, approve, or override an agent's actions before they take effect.
Prerequisites
What Is Human-in-the-Loop?
An autonomous agent can plan and act without a person reviewing every step, which is exactly what makes it useful — and exactly what makes a wrong decision risky. Human-in-the-loop means inserting a checkpoint where a person reviews, approves, or can override an agent's action before it actually takes effect, for the situations where the cost of a mistake is too high to let the agent proceed unsupervised.
Agent Decides Action
Not Yet ExecutedThe agent has decided, but hasn't acted yet.
Human Review Checkpoint
Blocks ExecutionA person reviews before anything takes effect.
Execute
Takes EffectRuns only after explicit approval.
Adjust
Sent BackThe agent revises its plan instead of acting.
Key Idea
Human-in-the-loop isn't about removing agent autonomy everywhere — it's about placing checkpoints where the risk of an unsupervised action is highest.
When to Use It
- Irreversible actions — sending money, deleting data, or sending a message on someone's behalf, where a mistake can't simply be undone.
- High-stakes decisions — actions with legal, financial, medical, or safety consequences.
- Low-confidence situations — when the agent itself signals uncertainty, or when a task is unusual enough to fall outside well-tested behavior.
- New or unproven agent behavior — while an agent workflow is still being validated in production, before trusting it to run fully autonomously.
Tip
Human-in-the-loop and prompt injection defense reinforce each other — requiring approval for consequential actions limits the damage a successful injection can do.
A Real-World Example
An agent that drafts and sends customer emails might be trusted to draft freely, but a team may still require a person to approve the draft before it actually sends — especially early on, while confidence in the agent's judgment is still being established. Over time, as the agent's draft quality proves reliable for a narrow, well-understood category of emails, the team might choose to let those specific cases send automatically while keeping human review for anything unusual.
Common Mistakes
Giving an agent full autonomy for irreversible actions
Actions that can't be undone deserve a human checkpoint, especially while an agent's reliability is still unproven.
Requiring approval for everything, including trivial actions
Over-applying human review defeats the purpose of automation and creates approval fatigue, making reviewers more likely to rubber-stamp requests without real scrutiny.
Treating human-in-the-loop as a permanent design instead of a phased one
As confidence in an agent's behavior grows for well-understood cases, the review threshold can often be relaxed for those specific cases while keeping it for higher-risk ones.
Not giving the reviewer enough context to make a real decision
A meaningful approval step needs to show what the agent is about to do and why, not just a bare confirm button.
Interview Question
What is human-in-the-loop, and when would you require it for an agent?
Human-in-the-loop means inserting a checkpoint where a person can review, approve, or override an agent's action before it takes effect. You'd require it for irreversible or high-stakes actions — sending money, deleting data, anything with legal or safety consequences — and for agent behavior that's new or still being validated in production. It's not about removing autonomy everywhere, it's about placing review where the cost of a mistake is highest, and relaxing it over time for narrow, well-understood cases once the agent has proven reliable there.
What an interviewer may ask next
- Why might requiring approval for every single agent action actually backfire?
- How does human-in-the-loop relate to defending against prompt injection?
- How would you decide when it's safe to relax a human review requirement?
Explain It in 30 Seconds
Human-in-the-loop inserts a checkpoint where a person reviews or approves an agent's action before it takes effect, for situations where a mistake would be costly or hard to undo — irreversible actions, high-stakes decisions, or unproven agent behavior. It's not about removing autonomy everywhere, just placing review where risk is highest, and it also limits the damage from a successful prompt injection by requiring approval before a consequential action executes.