Jailbreaks & Adversarial Inputs
A jailbreak is an input crafted specifically to bypass a model’s safety training or stated instructions.
Prerequisites
Overview
A jailbreak differs from prompt injection in intent: injection typically hijacks a model via untrusted third-party content, while a jailbreak is usually the end user themselves directly trying to get the model to ignore its own safety training or system instructions.
Where It Fits
Crafted Adversarial Input
Model
Safety training is a defense layer, not a guarantee.
Output Guardrail Check
Blocked or Allowed
Key Points
- Roleplay and hypothetical framing
- A common jailbreak pattern asks a model to adopt a persona or hypothetical framing that supposedly excuses otherwise-restricted output.
- Defense in depth
- Since no single defense against jailbreaks is complete, output-side guardrails and monitoring matter as much as the model’s own training.
- Evolving technique
- Jailbreak techniques change constantly as models are updated — a static defense list goes stale quickly, which is why ongoing red teaming matters.
Interview Question
How is a jailbreak different from a prompt injection attack?
Prompt injection typically comes from untrusted third-party content — a retrieved document or tool result — hijacking the model. A jailbreak is usually the end user themselves directly trying to get the model to bypass its own safety training or system instructions, often through roleplay or hypothetical framing. Both exploit the same underlying issue — a model can’t perfectly separate instructions from intent — but the source and framing differ.
Explain It in 30 Seconds
A jailbreak is an adversarial input, usually from the end user directly, designed to bypass a model’s safety training or system instructions — commonly through roleplay or hypothetical framing — and needs output-side guardrails and ongoing red teaming as a defense, since no single control fully prevents it.
Real-World Stack
Technologies commonly used to implement this in production.