ChatGPT-like Architecture
This architecture describes how a conversational AI product connects a chat interface, backend, and model provider together.
Prerequisites
Why Start Here?
A conversational AI product — a chat interface backed by a language model — is one of the simplest complete AI systems you can build, and almost every more complex architecture in this curriculum is a variation on it. Understanding how its pieces fit together is the foundation for reasoning about RAG systems, agents, and everything else. This lesson is deliberately vendor-neutral: it explains the pattern behind any chat-style AI product, not a specific company's implementation.
Key Idea
A chat product is not just "a UI that calls a model." Production versions add a layer of orchestration between the two to manage context, history, and safety.
The Request Flow
At its core, a single turn in a conversation moves through a small number of stages: the client sends a message, the backend assembles everything the model needs to know, the model generates a response, and that response is returned and stored as part of the conversation history.
Client (chat UI)
Message SentThe user submits one turn of the conversation.
API / Backend
Receives TurnThe application backend, not the model itself.
Context Assembly (system prompt + history)
Builds the PromptCombines system prompt, history, and retrieved content.
Model Provider
Generates ReplyProduces a response from the assembled context.
Response Handling
Stores TurnSaved as part of the conversation history.
Client (rendered reply)
Shown to UserWhat the user actually sees in the chat UI.
"Context assembly" is doing more work than it looks like: it combines the system prompt, as much conversation history as fits in the context window, and — in more advanced versions — retrieved documents or tool results. This is the step where context-window limits, token budgets, and prompt construction all become real engineering concerns rather than abstract ideas.
What the Backend Owns
- Conversation state
- Storing message history so a multi-turn conversation can be reconstructed on each new request — the model itself is stateless between calls.
- Context assembly
- Deciding what actually gets sent to the model on this turn: system prompt, trimmed or summarized history, and any retrieved or tool content.
- Streaming
- Relaying the model's output back to the client incrementally as it's generated, rather than waiting for the full response — this is what makes a chat UI feel responsive.
- Safety and validation
- Applying guardrails to user input and model output before either is trusted — the backend is the natural place to enforce this consistently.
Important
The model provider is stateless and vendor-specific; everything about conversation state, history management, and safety is the application's responsibility, not the model's.
Simple vs. Production
Client
Backend forwards message + history
Model
Return response
Client
Auth + rate limiting
Context assembly (history, retrieval, tools)
Guardrails on input
Model (via gateway)
Guardrails on output
Store + stream response
The added complexity in the production version isn't accidental — each stage solves a specific problem: auth and rate limiting protect the backend from abuse and runaway cost, guardrails reduce the risk of unsafe input or output, and a gateway (covered in its own lesson) centralizes how the model itself is called.
Failure Modes
- Model provider latency or outage — the user-facing request has no response until the backend gets one back, or times out and degrades gracefully.
- Context window overflow — a long conversation eventually needs truncation or summarization, or requests start failing.
- Streaming interruption — a dropped connection mid-stream needs a defined behavior, not a silently truncated reply.
- Unsafe input or output — without validation, a single bad turn can produce a harmful or off-policy response.
Common Mistakes
Sending the entire raw conversation history on every turn
Without trimming or summarization, history grows until it exceeds the context window or becomes unnecessarily expensive.
Treating the model call as instantaneous
No timeout or fallback around a model call means a slow or hanging provider directly hangs the user-facing request.
Coupling the backend tightly to one provider's API
Provider-specific logic spread throughout the application makes it hard to add a fallback provider or switch models later.
Skipping input/output validation because "it's just a chatbot"
Even a simple chat product can produce or accept unsafe content without basic guardrails.
Storing full conversation history with no retention policy
Conversation logs often contain sensitive user content and need the same data-handling discipline as any other user data.
Interview Question
How would you architect a production chat-based AI product, and what does the backend need to own beyond just calling the model?
I'd separate the model provider, which is stateless and only sees what it's given on each call, from the backend, which owns everything about conversation state: storing history, assembling context for each turn within the token budget, streaming the response back to the client, and applying safety checks on both input and output. A minimal version can just forward the message and history straight to the model, but a production version adds auth and rate limiting to control cost and abuse, guardrails around input and output, and typically a gateway layer to manage the model call itself — retries, fallback, and observability — rather than calling the provider directly from application code.
What an interviewer may ask next
- Why is the model provider considered stateless, and what does that mean for the backend?
- What would you do if a conversation grows too long for the context window?
- What should happen if the model provider is slow or unavailable mid-conversation?
- Why shouldn't application code call a model provider directly in a production system?
Explain It in 30 Seconds
A conversational AI product connects a chat client to a model provider through a backend that owns everything the model itself doesn't: conversation history, context assembly within the token budget, streaming, and safety checks on input and output. The model is stateless between calls, so all of that state and orchestration lives in the application. Production versions add auth, rate limiting, guardrails, and usually a gateway layer between the backend and the model provider.