AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Beginner30–45 min

AI Chat Application

Learn the architecture behind a conversational AI product — the pattern nearly every other project in this section builds on.

What You Will Build

A small, illustrative chat application architecture: a chat UI that sends messages to a backend, which assembles context, calls a model, and streams the response back. This is the same pattern behind any ChatGPT-like product, built here as a learning exercise, not a production clone.

Learning Objectives

  • Understand the request lifecycle of a conversational AI application

  • Distinguish system, user, and assistant messages

  • Understand why conversation history has to be managed, not just stored

  • Understand where streaming and error handling fit in

Prerequisites

Concepts Used

LLM
System Prompt
Context Window
Token
Streaming
ChatGPT-like Architecture

Architecture

User (Chat UI)

Sends a Message

The new turn in the conversation.

sent to

Application Backend

Owns Conversation State

Coordinates the whole request lifecycle.

triggers

Context Assembly

Shared Budget

System prompt, trimmed history, and new message compete for it.

sent to

LLM

Stateless Call

Generates a new assistant message from the full context.

generates

Streamed Response

Token by Token

Improves perceived responsiveness for longer answers.

relayed to

Chat UI

Renders Incrementally

Shows the response as it arrives.

AI chat application request flow

Step 1 — Define the Message Model

What are we doing? Defining the shape of a single conversation message. Why? Every chat-based LLM API works with a list of role-tagged messages — system, user, assistant — not a single blob of text. How it works: each message has a role and content; the full conversation is an ordered array of these.

message_model.py (illustrative)
class Message:
    def __init__(self, role: str, content: str):
        self.role = role        # "system" | "user" | "assistant"
        self.content = content

conversation = [
    Message("system", "You are a concise, helpful assistant."),
    Message("user", "What is a vector database?"),
]

Step 2 — Assemble Context for Each Turn

What are we doing? Deciding exactly what gets sent to the model on a given turn. Why? The context window is a shared, limited budget — the system prompt, as much history as reasonably fits, and the new user message all compete for it. How it works: trim or summarize older messages before they exceed the budget, rather than sending the full history indefinitely.

Key Idea

Context assembly is application logic, not model logic — the model just sees whatever sequence of messages it's given.

Step 3 — Call the Model

What are we doing? Sending the assembled conversation to the model and getting a response. Why? This is the step that actually produces new content. How it works: the request includes the full message list; the model generates a new assistant message conditioned on all of it.

call_model.py (illustrative pseudocode)
def get_response(conversation, timeout=10):
    try:
        response = model.generate(messages=conversation, timeout=timeout)
        return Message("assistant", response.text)
    except ModelTimeout:
        return Message("assistant", "Sorry, that took too long — please try again.")

Step 4 — Stream the Response

What are we doing? Sending tokens to the UI as they're generated, instead of waiting for the full response. Why? A decoder already generates one token at a time — streaming just exposes that to the user, dramatically improving perceived responsiveness for longer answers. For the full mechanics — cancellation, backpressure, partial-response errors — see the dedicated Streaming AI project.

LLM Generates Token

Already Sequential

A decoder produces one token at a time regardless.

streamed to

Send to Client

As It's Generated

No waiting for the full response.

displayed via

Render Incrementally

Perceived Speed

The user sees progress immediately.

then

Repeat Until Done

Until Complete

Continues until the model finishes generating.

Step 5 — Handle Failure

What are we doing? Deciding what happens when something goes wrong. Why? AI systems fail differently from typical backends — a slow provider, a timeout, or a malformed response are routine, not exceptional. How it works: wrap the model call with a timeout and a clear fallback message, and never let a hung request block the UI indefinitely.

  • Sending unbounded conversation history

    Without trimming, a long conversation eventually exceeds the context window or gets unnecessarily expensive.

  • No timeout around the model call

    A slow provider directly hangs the user-facing request without one.

  • Treating the model as deterministic

    The same input can produce a different output on different calls — design for that, don't assume identical responses.

  • Coupling application code tightly to one provider

    Provider-specific logic spread through the codebase makes it hard to add a fallback or switch providers later.

  • Skipping input/output validation because "it's just a chatbot"

    Even a simple chat product benefits from basic guardrails on both sides of the model call.

Challenges

Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.

Challenge 1: Add conversation summarization

When history grows too long, summarize older messages instead of dropping them outright.

Hint

Summarize in a separate model call, then replace the oldest messages with the summary.

Challenge 2: Add a system-prompt-driven persona

Change the system prompt to give the assistant a distinct role, and observe how consistently the model follows it.

Challenge 3: Handle a mid-stream failure

Simulate the model failing partway through a streamed response, and design what the UI should show.

Expected approach

Show the partial response with a clear "generation interrupted" indicator rather than silently truncating it.

Design Review

Before moving on, think through these questions the way a reviewer would.

  • Where does conversation state actually live, and who owns it?

  • What happens to this application if the model provider is slow for 30 seconds?

  • Where would you add a gateway if this needed to support multiple model providers?

  • How would you prevent a single very long conversation from breaking the app?

Interview Questions

Walk me through what happens when a user sends a message in a chat application like this.

The client sends the new message to the backend, which assembles the full context — system prompt, trimmed history, and the new message — and calls the model. The model generates a response, typically streamed back token by token, which the backend relays to the client as it arrives. Once complete, the new exchange is added to stored conversation history for the next turn.

  • Model is stateless between calls
  • Backend owns context assembly and history
  • Streaming improves perceived latency, not total time

How would you handle a very long conversation that starts approaching the context window limit?

I'd trim or summarize older messages rather than sending the full history indefinitely — keeping the system prompt and recent messages intact, and either dropping or summarizing the oldest content once the running total approaches the budget.

  • Context window is a shared budget across prompt, history, and output
  • Summarization preserves more information than simple truncation

Why shouldn't application code call the model provider directly from many places?

Doing so scatters authentication, retry, and error-handling logic across the codebase, and makes it hard to add a fallback provider or switch models later. Centralizing that behind a gateway keeps it consistent — covered in depth in the LLM Gateway project.

  • Provider abstraction
  • Consistent retry/fallback behavior
  • Easier to evolve later

Explain It in 30 Seconds

This project builds the architecture behind a conversational AI product: a chat UI sends messages to a backend, which assembles context from the system prompt and trimmed history, calls the model, and streams the response back token by token. The model itself is stateless — the backend owns conversation state, context assembly, and error handling. It's the foundational pattern nearly every other AI application in this section builds on.

On this page