AI Chat Application
Learn the architecture behind a conversational AI product — the pattern nearly every other project in this section builds on.
What You Will Build
A small, illustrative chat application architecture: a chat UI that sends messages to a backend, which assembles context, calls a model, and streams the response back. This is the same pattern behind any ChatGPT-like product, built here as a learning exercise, not a production clone.
Learning Objectives
Understand the request lifecycle of a conversational AI application
Distinguish system, user, and assistant messages
Understand why conversation history has to be managed, not just stored
Understand where streaming and error handling fit in
Prerequisites
Concepts Used
Architecture
User (Chat UI)
Sends a MessageThe new turn in the conversation.
Application Backend
Owns Conversation StateCoordinates the whole request lifecycle.
Context Assembly
Shared BudgetSystem prompt, trimmed history, and new message compete for it.
LLM
Stateless CallGenerates a new assistant message from the full context.
Streamed Response
Token by TokenImproves perceived responsiveness for longer answers.
Chat UI
Renders IncrementallyShows the response as it arrives.
Step 1 — Define the Message Model
What are we doing? Defining the shape of a single conversation message. Why? Every chat-based LLM API works with a list of role-tagged messages — system, user, assistant — not a single blob of text. How it works: each message has a role and content; the full conversation is an ordered array of these.
class Message:
def __init__(self, role: str, content: str):
self.role = role # "system" | "user" | "assistant"
self.content = content
conversation = [
Message("system", "You are a concise, helpful assistant."),
Message("user", "What is a vector database?"),
]Step 2 — Assemble Context for Each Turn
What are we doing? Deciding exactly what gets sent to the model on a given turn. Why? The context window is a shared, limited budget — the system prompt, as much history as reasonably fits, and the new user message all compete for it. How it works: trim or summarize older messages before they exceed the budget, rather than sending the full history indefinitely.
Key Idea
Context assembly is application logic, not model logic — the model just sees whatever sequence of messages it's given.
Step 3 — Call the Model
What are we doing? Sending the assembled conversation to the model and getting a response. Why? This is the step that actually produces new content. How it works: the request includes the full message list; the model generates a new assistant message conditioned on all of it.
def get_response(conversation, timeout=10):
try:
response = model.generate(messages=conversation, timeout=timeout)
return Message("assistant", response.text)
except ModelTimeout:
return Message("assistant", "Sorry, that took too long — please try again.")Step 4 — Stream the Response
What are we doing? Sending tokens to the UI as they're generated, instead of waiting for the full response. Why? A decoder already generates one token at a time — streaming just exposes that to the user, dramatically improving perceived responsiveness for longer answers. For the full mechanics — cancellation, backpressure, partial-response errors — see the dedicated Streaming AI project.
LLM Generates Token
Already SequentialA decoder produces one token at a time regardless.
Send to Client
As It's GeneratedNo waiting for the full response.
Render Incrementally
Perceived SpeedThe user sees progress immediately.
Repeat Until Done
Until CompleteContinues until the model finishes generating.
Step 5 — Handle Failure
What are we doing? Deciding what happens when something goes wrong. Why? AI systems fail differently from typical backends — a slow provider, a timeout, or a malformed response are routine, not exceptional. How it works: wrap the model call with a timeout and a clear fallback message, and never let a hung request block the UI indefinitely.
Sending unbounded conversation history
Without trimming, a long conversation eventually exceeds the context window or gets unnecessarily expensive.
No timeout around the model call
A slow provider directly hangs the user-facing request without one.
Treating the model as deterministic
The same input can produce a different output on different calls — design for that, don't assume identical responses.
Coupling application code tightly to one provider
Provider-specific logic spread through the codebase makes it hard to add a fallback or switch providers later.
Skipping input/output validation because "it's just a chatbot"
Even a simple chat product benefits from basic guardrails on both sides of the model call.
Challenges
Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.
Challenge 1: Add conversation summarization
When history grows too long, summarize older messages instead of dropping them outright.
Hint
Summarize in a separate model call, then replace the oldest messages with the summary.
Challenge 2: Add a system-prompt-driven persona
Change the system prompt to give the assistant a distinct role, and observe how consistently the model follows it.
Challenge 3: Handle a mid-stream failure
Simulate the model failing partway through a streamed response, and design what the UI should show.
Expected approach
Show the partial response with a clear "generation interrupted" indicator rather than silently truncating it.
Design Review
Before moving on, think through these questions the way a reviewer would.
Where does conversation state actually live, and who owns it?
What happens to this application if the model provider is slow for 30 seconds?
Where would you add a gateway if this needed to support multiple model providers?
How would you prevent a single very long conversation from breaking the app?
Interview Questions
Walk me through what happens when a user sends a message in a chat application like this.
The client sends the new message to the backend, which assembles the full context — system prompt, trimmed history, and the new message — and calls the model. The model generates a response, typically streamed back token by token, which the backend relays to the client as it arrives. Once complete, the new exchange is added to stored conversation history for the next turn.
- Model is stateless between calls
- Backend owns context assembly and history
- Streaming improves perceived latency, not total time
How would you handle a very long conversation that starts approaching the context window limit?
I'd trim or summarize older messages rather than sending the full history indefinitely — keeping the system prompt and recent messages intact, and either dropping or summarizing the oldest content once the running total approaches the budget.
- Context window is a shared budget across prompt, history, and output
- Summarization preserves more information than simple truncation
Why shouldn't application code call the model provider directly from many places?
Doing so scatters authentication, retry, and error-handling logic across the codebase, and makes it hard to add a fallback provider or switch models later. Centralizing that behind a gateway keeps it consistent — covered in depth in the LLM Gateway project.
- Provider abstraction
- Consistent retry/fallback behavior
- Easier to evolve later
Explain It in 30 Seconds
This project builds the architecture behind a conversational AI product: a chat UI sends messages to a backend, which assembles context from the system prompt and trimmed history, calls the model, and streams the response back token by token. The model itself is stateless — the backend owns conversation state, context assembly, and error handling. It's the foundational pattern nearly every other AI application in this section builds on.