Guardrails
Guardrails are checks that constrain model input or output to prevent unsafe, incorrect, or off-policy behavior.
Prerequisites
What Are Guardrails?
A model's behavior is shaped by its prompt and training, but neither guarantees it will always stay within acceptable bounds. Guardrails are explicit checks — separate from the model itself — that sit around it, validating input before it reaches the model and validating output before it's used, to catch cases the model's own judgment might miss.
Input
UncheckedNot yet validated against any policy.
Input Guardrail
Checks Before ModelValidates input before it reaches the model.
Model
Model's Own JudgmentNot guaranteed to stay within bounds on its own.
Output Guardrail
Checks Before UseValidates output before it's used.
Final Result
Passed Both ChecksOnly reaches the user after both guardrails.
Key Idea
Guardrails don't rely on the model to police itself — they're an external check the system enforces regardless of what the model decides to do.
Types of Guardrails
- Input validation — checking incoming requests for disallowed content, injection attempts, or malformed data before they reach the model.
- Output validation — checking a model's response against expected format, allowed content, or policy before it's shown to a user or acted on.
- Tool-call guardrails — for agents, validating that a requested tool call is within allowed scope and arguments before it actually executes.
- Rate and scope limits — bounding how much an agent or feature can do in a given time window, limiting the damage from a misbehaving loop.
Tip
Guardrails work well alongside human-in-the-loop review — guardrails catch clear-cut violations automatically, while human review handles ambiguous or high-stakes cases.
A Real-World Example
An agent with access to a "send email" tool might have a guardrail that blocks any tool call attempting to send to an external domain outside an approved list, regardless of what the model decided to do. Even if a prompt injection or a model mistake led the agent to attempt this, the guardrail — enforced outside the model's control — stops the action before it takes effect.
Common Mistakes
Relying on the model to enforce its own boundaries
Instructions in a prompt influence behavior but don't guarantee it — guardrails need to be enforced outside the model's control.
Only validating input, not output
A model can still produce unsafe or incorrect output from a perfectly reasonable input — both directions need checks.
Making guardrails so strict they block legitimate use
Overly aggressive guardrails create false positives that frustrate real users — they need to be tuned, not just tightened indefinitely.
Not logging when a guardrail triggers
Guardrail activity is valuable signal for understanding misuse patterns and tuning the system — silently blocking without logging wastes that signal.
Interview Question
What are guardrails, and why can't you rely on the model alone to stay within acceptable bounds?
Guardrails are explicit checks outside the model itself that validate input before it reaches the model and validate output — or a proposed tool call — before it's used or executed. You can't rely on the model alone because prompt instructions strongly influence behavior but don't guarantee it, especially under adversarial input like prompt injection. In practice that means input validation, output validation, and for agents specifically, checking that a tool call is within allowed scope before it actually runs — enforced by the surrounding system, not left to the model's own judgment.
What an interviewer may ask next
- Why aren't prompt instructions alone sufficient to keep a model within bounds?
- How do guardrails complement human-in-the-loop review?
- What could go wrong if guardrails only checked input and never output?
Explain It in 30 Seconds
Guardrails are checks outside the model itself — validating input before it reaches the model and validating output or proposed tool calls before they're used — because prompt instructions influence behavior but don't guarantee it, especially under adversarial input. They work alongside human-in-the-loop review: guardrails catch clear-cut violations automatically, humans handle ambiguous or high-stakes cases. Both input and output need checks, since a model can still produce unsafe output from reasonable input.