AI Agent
Apply the agent lessons to a bounded, tool-using agent loop with retries, memory, and a human approval checkpoint.
What You Will Build
An illustrative agent that pursues a goal across multiple steps: it plans, calls tools, incorporates results, and decides when to stop — with explicit bounds, memory, and a human approval checkpoint for a risky action. This project applies the AI Agent, Agent Loop, Planning, Tool Use, AI Memory, and Human-in-the-Loop lessons.
Learning Objectives
Implement a bounded agent loop with a maximum step count
Apply tool calling with schema validation and failure handling
Add a human approval checkpoint for a consequential action
Understand where an agent can fail and how to observe it
Prerequisites
Concepts Used
Architecture
User Goal
Starting PointThe high-level task the agent pursues.
Agent Runtime
Owns the LoopEnforces bounds independent of the model.
Planner
One Step at a TimeDecides just the next action given what happened so far.
Memory
Retains StateKeeps track of history across steps.
Tool Registry + Executor
Validated ExecutionChecks calls against schemas before running them.
Human Approval (if needed)
Risk CheckpointRequired before an irreversible action executes.
Final Answer
Loop ExitReturned once the agent decides it's done.
Step 1 — Accept a Goal and Plan
What are we doing? Taking a high-level goal and breaking it into steps. Why? Some goals can't be handled in one model call. How it works: this project uses dynamic planning — deciding just the next step at a time based on what's happened so far, adjusting as new information appears.
Goal
High-Level TaskWhat the agent is trying to accomplish.
Observe State
Current ContextWhat has happened so far in this run.
Decide Next Action
Dynamic PlanningChooses just the next step, not the whole plan.
Act
Executes a ToolCarries out the chosen action.
Observe Result
New InformationWhat the action actually returned.
Repeat or Finish
Loop DecisionContinue the cycle, or return a final answer.
Step 2 — Define the Tool Registry
What are we doing? Defining which tools the agent can call and validating requested calls against their schemas. Why? An agent should only have access to the tools its task actually needs, following least privilege — and every call needs validation before it executes.
TOOLS = {
"search_docs": {"schema": search_docs_schema, "handler": search_docs},
"send_email": {"schema": send_email_schema, "handler": send_email, "requires_approval": True},
}
def execute_tool(name, arguments):
if name not in TOOLS:
raise UnknownTool(name)
validate(arguments, TOOLS[name]["schema"])
if TOOLS[name].get("requires_approval"):
require_human_approval(name, arguments)
return TOOLS[name]["handler"](**arguments)Step 3 — Implement the Bounded Loop
What are we doing? Running the observe-decide-act cycle with an explicit step limit. Why? Without a hard limit, a confused agent can loop far longer than intended, accumulating cost.
def run_agent(goal, max_steps=8):
state = {"goal": goal, "history": []}
for step in range(max_steps):
decision = plan_next_step(state)
if decision.type == "final_answer":
return decision.answer
result = execute_tool(decision.tool, decision.arguments)
state["history"].append((decision, result))
return "Stopped: reached maximum steps without a final answer."Step 4 — Add Memory
What are we doing? Retaining relevant information across steps so the agent doesn't treat each step as isolated. Why? Without memory, the agent can repeat work or lose track of earlier findings. How it works: the running `history` in the loop above is the simplest form of memory — for longer-running or multi-session agents, this could be summarized or persisted.
Step 5 — Require Approval for Risky Actions
What are we doing? Adding a checkpoint before an irreversible action executes. Why? An agent's judgment can be wrong, and some mistakes — sending a message, spending money — can't simply be undone.
Important
Tool results should be treated as untrusted input, the same as any retrieved document — a tool call to something external can return content that attempts to redirect the agent.
Step 6 — Handle Failure Modes
No maximum step limit
A stuck or confused agent can loop far longer than intended without a hard ceiling.
Granting broad tool access by default
An agent should have only the tools its specific task requires.
Treating tool results as trusted by default
Content returned from a tool can carry injected instructions.
No retry or recovery path for a failed tool call
A single tool error shouldn't necessarily end the whole run.
No observability into agent decisions
Without logging each step, debugging why an agent did something is very difficult.
Challenges
Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.
Challenge 1: Add retry handling
Retry a failed tool call once with adjusted arguments before giving up.
Challenge 2: Add human approval
Require explicit approval before any tool marked as requiring it actually executes.
Challenge 3: Prevent infinite loops
Detect when the agent is repeating the same action without progress, and stop early.
Challenge 4: Add reflection
After each tool result, have the agent briefly assess whether it actually made progress before continuing.
Design Review
Before moving on, think through these questions the way a reviewer would.
What is the maximum damage this agent could do if it misused its most powerful tool?
How would you observe what this agent actually did during a run, after the fact?
Where would you add a human checkpoint if this agent handled financial transactions?
What would you change if this agent needed to coordinate with a second, specialized agent?
Interview Questions
How would you design a production agent system, and what keeps it safe and bounded?
I'd separate the agent loop from the runtime around it. The runtime owns a tool registry that validates and executes calls, memory for relevant state, and enforces limits independent of the model's own decisions — a maximum step count, scoped tool access following least privilege, and required human approval for irreversible actions. I'd treat tool results as untrusted input, the same as retrieved documents, since they can carry injected instructions. I'd start with a single well-scoped agent rather than multi-agent orchestration, unless the task genuinely spans distinct specialties.
- Runtime enforces bounds, not the model
- Least privilege on tool access
- Human approval for irreversible actions
- Tool results are untrusted input
What would you do if an agent got stuck calling the same tool repeatedly without making progress?
I'd add a check that detects repeated identical or near-identical actions without a changing result, and terminate the run early with a clear message rather than letting it burn through the full step budget. I'd also log the sequence so it's possible to see why the agent got stuck.
- Detect lack of progress, not just step count
- Fail with a clear message
- Log for debugging
Why treat a tool's result as untrusted input?
Because a tool can return content from outside the application — a web page, a document, an API response — and the model has no inherent way to distinguish that content from a legitimate instruction. This is the same trust-boundary problem as prompt injection, just occurring inside the agent loop instead of at the initial prompt.
- Same root cause as prompt injection
- Applies to any external tool result
Explain It in 30 Seconds
This project applies the agent lessons to a bounded, tool-using agent: it plans one step at a time, calls tools through a validated registry, retains relevant state as memory, and requires human approval before irreversible actions. The runtime — not the model — enforces a maximum step count and scoped tool permissions, and tool results are treated as untrusted input, the same trust-boundary concern as prompt injection.