AI Security Architecture
AI security architecture designs how an AI system enforces authentication, data protection, and safe tool use end to end.
Prerequisites
The New Trust Boundary
Traditional application security draws a fairly clear line: user input is untrusted, application code and its own data are trusted. AI systems blur that line, because a model treats everything in its context window as just text — a retrieved document, a tool result, or a piece of conversation history can all carry instructions the model may act on, whether or not the application intended them as instructions.
User Input
Input BoundaryValidated before it reaches the model.
Application
Orchestrates the CallFetches context and applies boundary checks.
Untrusted Retrieved Content / Tool Results
Retrieval BoundaryData to reason about, not instructions to follow.
Model Context
Everything Is TextThe model can't tell instructions from data here.
Model Output
Output BoundaryValidated before it's shown, stored, or acted on.
Important
This lesson builds directly on the Prompt Injection and Guardrails lessons — the point here is architectural: where in the system these defenses actually need to live.
Security Boundaries in an AI System
- Input boundary
- Where user input first enters the system — validated before it reaches the model or gets embedded into a prompt.
- Retrieval boundary
- Where retrieved content enters the model's context — this content should be treated as data to reason about, not instructions to follow, and access-filtered by permission metadata.
- Tool boundary
- Where a model-requested action actually executes — this is where least privilege and approval checkpoints matter most, since it's the point where a decision becomes a real-world effect.
- Output boundary
- Where the model's response leaves the system — validated before it's shown to a user, stored, or used to trigger another action.
Key Idea
Each boundary needs its own check. A strong input validator does nothing to protect the tool boundary, and vice versa.
Code Example
Illustrative pseudocode showing where checks belong relative to the model call, not a complete security implementation:
def handle_request(user_input, tenant_id):
validate_input(user_input) # input boundary
retrieved = retrieve(user_input, tenant_filter=tenant_id) # retrieval boundary:
context = mark_as_untrusted_data(retrieved) # tag, don't treat as instructions
response = model.generate(system_prompt, context, user_input)
if response.requested_tool_call:
if not tool_is_allowed(response.tool_name, tenant_id): # tool boundary
raise PermissionDenied()
if requires_approval(response.tool_name):
await human_approval(response)
result = execute_tool(response.tool_name, response.arguments)
return validate_output(response) # output boundaryLeast Privilege in Practice
- Scope tool access per use case — an agent handling customer questions shouldn't have the same tool access as one handling account administration.
- Scope data access per request — retrieval should be filtered by the requesting user's or tenant's actual permissions, not the broadest access available to the system.
- Separate read and write capabilities — a tool that can read data is a much smaller risk surface than one that can also modify or send something.
- Keep secrets out of prompts entirely — API keys, credentials, and internal identifiers shouldn't be constructed into a prompt where a model could echo them back.
Common Mistakes
Treating a strong system prompt as a security control
Prompt instructions influence behavior but are not a reliable security boundary against adversarial input.
Trusting retrieved content or tool results by default
Anything the model reads that wasn't written by the application itself should be treated as untrusted data.
Granting broad tool permissions for convenience
Excess permissions turn any single injection or mistake into a much larger incident than it needs to be.
Validating only input and not output
A model can produce unsafe or policy-violating output even from a reasonable, fully validated input.
Logging full prompts and responses without redaction
Logs are a common place for sensitive user data or secrets to leak if captured indiscriminately.
No audit trail for tool calls or data access
Without logging what an agent accessed or executed and why, incident response and compliance both become far harder.
Interview Question
How would you design the security architecture for a system where an LLM can read retrieved documents and call tools?
I'd design around four boundaries rather than one: where user input enters, where retrieved content enters the model's context, where a tool call actually executes, and where the model's output leaves the system — each needs its own check, since a strong input validator does nothing to protect the tool boundary. Retrieved content and tool results get treated as untrusted data, not instructions, building on the same trust-boundary problem as prompt injection. At the tool boundary specifically, I'd enforce least privilege — scoping tool and data access to what the specific use case actually needs — and require human approval for anything irreversible. I'd validate output before it's used or shown, and keep an audit trail of tool calls and data access, since a system prompt alone is an influence on behavior, not a security guarantee.
What an interviewer may ask next
- Why isn't a well-written system prompt sufficient as a security boundary?
- Why do retrieved documents and tool results need to be treated as untrusted, even though they came from your own systems?
- What is the difference between the input boundary and the output boundary, and why do both need separate checks?
- How would you apply least privilege to an agent's tool access?
Explain It in 30 Seconds
AI security architecture treats a model's context window as a place where instructions and data blur together, so it defines separate checks at four boundaries: input, retrieval, tool execution, and output. Retrieved content and tool results are treated as untrusted data, not instructions — the same trust-boundary problem as prompt injection, addressed architecturally. Least privilege at the tool boundary and human approval for irreversible actions matter most, since that's where a decision becomes a real effect, and none of this replaces the need for an audit trail of what was accessed and executed.