Agent Architecture
Agent architecture describes how planning, memory, tools, and orchestration fit together in a production agent system.
Prerequisites
From a Loop to a System
The agent loop describes the repeating cycle a single agent runs through. Agent architecture is the surrounding system that makes that loop safe, observable, and bounded in production — the runtime, tool registry, memory store, and guardrails that sit around the loop itself.
User Goal
Starting PointThe task the user wants the agent to accomplish.
Agent Runtime
Surrounding SystemMakes the loop safe, observable, and bounded.
Planner
Breaks Down GoalSplits the goal into steps, upfront or dynamically.
Memory
Retains StateKeeps relevant information across steps or sessions.
Tool Registry + Executor
Runs ToolsValidates a requested call against its schema and executes it.
Guardrails
Enforces LimitsBlocks disallowed tools, bad arguments, or actions needing approval.
Final Response
Bounded OutputWhat the user sees once the loop finishes.
Key Idea
The runtime, not the model, is what enforces limits: maximum steps, allowed tools, timeouts, and approval checkpoints all live in the surrounding system, not in the model's own judgment.
What Each Piece Owns
- Planner
- Breaks the goal into steps, upfront or dynamically — see the Planning lesson for the tradeoffs between the two.
- Memory
- Retains relevant information across steps or sessions — could be as simple as the current run's state, or persistent across sessions.
- Tool registry and executor
- Defines which tools exist, validates a requested call against its schema, executes it, and returns the result — see the Tool Use lesson for how the agent selects among them.
- Guardrails
- Enforces boundaries independent of the model's own decisions — which tools are allowed, what arguments are valid, what actions need human approval.
- Evaluator
- Assesses whether the agent is making progress or should stop, escalate, or retry — related to Reflection, but implemented as part of the surrounding system rather than left entirely to the model.
Bounded Execution
An agent loop without limits can run far longer than intended, repeat unproductive actions, or take an action nobody actually wants executed. Production agent architecture bounds execution explicitly:
- Maximum steps — a hard ceiling on how many loop iterations a single run can take.
- Tool allowlists and scoped permissions — an agent should only have access to the tools its task actually requires, following least privilege.
- Human-in-the-loop checkpoints — required approval before irreversible or high-stakes actions, as covered in its own lesson.
- Timeouts — both per tool call and for the run as a whole, so a stuck agent doesn't run indefinitely.
Important
Guardrails and human-in-the-loop checkpoints aren't optional polish on an agent architecture — they're what makes autonomous execution acceptable to run against anything that matters.
Failure Modes
- Runaway loop — no step limit means a confused agent keeps going indefinitely, accumulating cost.
- Tool failure — a tool call errors or times out; the agent needs a defined path to recognize and recover, not silently continue on bad data.
- Prompt injection via tool results — content returned from a tool (a web page, a document) can attempt to redirect the agent; this is the same trust-boundary problem covered in the Prompt Injection lesson, now inside the agent loop itself.
- Excessive permissions — an agent with broad, unnecessary tool access turns any single mistake or injection into a much bigger problem.
- Silent failure — an agent that fails to make progress but reports success gives false confidence unless the evaluator or observability catches it.
Single Agent vs. Multi-Agent
One planner
One toolset
Simpler to reason about
Can become overloaded for very broad tasks
Specialist agents
Coordinator routes tasks
Better separation of concerns
More coordination overhead
A single, well-scoped agent is usually simpler to build, debug, and secure. Multi-agent orchestration — covered in the Agent Orchestration lesson — is worth the added coordination cost once a task genuinely spans multiple distinct specialties that don't fit well in one agent's toolset and instructions.
Common Mistakes
No maximum step limit
A stuck or confused agent can loop far longer than intended without a hard ceiling.
Granting broad tool access by default
An agent should have only the tools its specific task requires — not every tool the system happens to expose.
Treating tool results as trusted by default
Content returned from a tool can contain injected instructions and should be treated with the same caution as any other untrusted input.
No human checkpoint for irreversible actions
Sending messages, spending money, or deleting data deserves review, especially while an agent's reliability is unproven.
No observability into agent decisions
Without logging each step's reasoning, tool calls, and results, debugging why an agent did something is very difficult.
Reaching for multi-agent orchestration before it's needed
A single well-scoped agent is simpler to build and secure — orchestration adds coordination overhead that should be justified by a genuine multi-specialty need.
Interview Question
How would you architect a production agent system, and what would you put in place to keep it safe and bounded?
I'd separate the agent loop itself from the surrounding runtime that makes it safe to run in production. The runtime owns a tool registry that validates and executes tool calls, memory for relevant state, and guardrails that enforce limits independent of the model's own decisions — which tools are allowed, what arguments are valid, and which actions need human approval before they execute. I'd bound execution explicitly with a maximum step count and per-call timeouts, scope tool access to the minimum the task needs, and require a human checkpoint for anything irreversible. I'd also treat tool results as untrusted input, since content coming back from a tool can carry injected instructions the same way a retrieved document can. I'd start with a single well-scoped agent and only move to multi-agent orchestration if the task genuinely spans specialties that don't fit one agent's toolset.
What an interviewer may ask next
- Why does the runtime need to enforce limits rather than relying on the model to stop itself?
- Why should tool results be treated as untrusted input?
- When would you introduce multi-agent orchestration instead of a single agent?
- What would you log to make an agent's behavior debuggable in production?
Explain It in 30 Seconds
Agent architecture is the runtime around the agent loop that makes it safe in production: a planner, memory, a tool registry and executor, guardrails, and an evaluator. The runtime — not the model — enforces bounds like a maximum step count, scoped tool permissions, timeouts, and human approval checkpoints for irreversible actions. Tool results should be treated as untrusted input, since they can carry injected instructions. A single well-scoped agent is usually simpler than multi-agent orchestration, which is worth the added coordination cost only once a task genuinely spans multiple specialties.