AI Security
AI security protects AI systems from risks such as prompt injection, data leakage, and misuse of tools.
Prerequisites
Design vs. Implementation
AI Security Architecture explains the system-level design: where trust boundaries sit, and which parts of a system need their own checks. This lesson is about the concrete engineering controls that actually implement that design — the specific validation, authorization, and handling logic a team writes.
Key Idea
This lesson builds directly on Prompt Injection and Guardrails — the concern here is what code you actually write to enforce those defenses.
Core Engineering Controls
- Input validation
- Checking incoming requests — format, length, disallowed content — before they reach the model, catching obviously malformed or malicious input early.
- Output validation
- Checking a model's response — and any tool call it requests — against expected format and policy before it's used or shown.
- Authorization and least privilege
- Ensuring a request only accesses data and tools the requesting user or agent is actually permitted to use — enforced by application code, not by asking the model nicely.
- Secrets handling
- Keeping API keys and credentials out of prompts entirely — they belong in the application or gateway layer, never constructed into text a model reads.
- Data isolation
- Filtering retrieval and storage by tenant or user, enforced at the query level — the same concern covered architecturally in multi-tenant AI architecture.
- Audit logging
- Recording what was accessed, what tools were called, and by whom, so an incident can actually be investigated after the fact.
Code Example
Illustrative pseudocode for a tool-call authorization check:
def authorize_tool_call(user, tool_name, arguments):
if tool_name not in ALLOWED_TOOLS_FOR_ROLE[user.role]:
raise PermissionDenied(f"{user.role} cannot use {tool_name}")
if tool_name in SENSITIVE_TOOLS:
require_human_approval(user, tool_name, arguments)
audit_log(user, tool_name, arguments)
return TrueCommon Mistakes
Relying on the system prompt as the only defense
Prompt instructions influence behavior but are not a reliable enforcement mechanism against adversarial or malformed input.
Authorizing at the wrong layer
Checking permissions only in the UI, and not again in the backend or at the tool-execution layer, leaves a gap an attacker can bypass.
Embedding secrets into prompts or tool schemas
Any credential a model can read is a credential that could leak through model output.
No audit trail for tool calls or data access
Without one, investigating what actually happened during an incident becomes guesswork.
Validating input but not output
A model can produce unsafe or policy-violating output even from a fully validated, reasonable input.
Treating security as a one-time setup rather than continuous enforcement
New tools, new data sources, and new agent capabilities each need their own security review as the system evolves.
Interview Question
What engineering controls would you actually implement to secure an AI application that uses tools and retrieval?
I'd implement validation on both sides of the model call — checking input before it reaches the model, and checking output, including any requested tool call, before it's used or executed. Authorization needs to be enforced in application code at the point of data access or tool execution, not assumed from the system prompt, following least privilege so a request only touches what it's actually permitted to. Secrets stay entirely out of prompts, living in the application or gateway layer instead. For multi-tenant systems, data access needs to be filtered by tenant at the query level, and every tool call and data access should be audit-logged so an incident can actually be investigated afterward. None of this is a one-time setup — new tools and data sources each need their own review.
What an interviewer may ask next
- Why is checking permissions only in the UI not sufficient?
- Why can't a system prompt be relied on as a security control?
- What would you log to make a security incident actually investigable?
Explain It in 30 Seconds
AI security engineering means implementing the concrete controls behind a system's security design: input and output validation, authorization enforced in application code rather than assumed from the system prompt, secrets kept entirely out of prompts, tenant-filtered data access, and audit logging for tool calls and data access. It builds directly on prompt injection and guardrails — the focus here is the actual code that enforces those defenses, applied continuously as new tools and data sources are added, not set up once.