AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Intermediate30–45 min

Streaming AI Application

Apply the Streaming lesson to the full mechanics: time to first token, cancellation, backpressure, and mid-stream errors.

What You Will Build

An illustrative streaming pipeline that sends model output incrementally to a client, handles a client disconnecting mid-stream, and handles a failure partway through generation. No real LLM provider is required — the streaming mechanics are illustrated directly.

Learning Objectives

  • Understand why streaming improves perceived latency without reducing total generation time

  • Implement cancellation when a client disconnects mid-stream

  • Handle an error that occurs after some tokens have already been sent

  • Understand backpressure and why a client can't always keep up

Prerequisites

Concepts Used

Streaming
Decoder
LLM Gateway

Architecture

Request

Starts Generation

Kicks off the model call.

triggers

LLM Generates Token

Same Underlying Rate

Streaming doesn't speed this up.

streamed to

Send to Client

Immediately

As soon as each token exists.

displayed via

Client Renders

Time to First Token

What actually improves perceived speed.

then

Repeat Until Done or Cancelled

Loop Exit

Stops on completion or client disconnect.

Streaming request

Step 1 — Compare Normal vs. Streaming

What are we doing? Contrasting a request that waits for the full response with one that streams incrementally. Why? This makes the actual improvement — time to first token, not total time — concrete before adding the more complex handling.

Normal (request/response)

Request

Wait

Complete Response

Streaming

Request

Token 1

Token 2

Token 3

…

Complete

Step 2 — Implement a Basic Stream

What are we doing? Sending each generated token to the client as soon as it exists. How it works: illustrated here as a generator that yields tokens one at a time, standing in for a real model's token-by-token output.

basic_stream.py (illustrative pseudocode)
def stream_response(prompt):
    for token in model.generate_stream(prompt):
        yield token  # sent to the client immediately

# Client side: render each token as it arrives
for token in stream_response(prompt):
    append_to_ui(token)

Step 3 — Handle Cancellation

What are we doing? Stopping generation when the client disconnects or the user navigates away. Why? Without this, the backend keeps generating and consuming provider resources for a response nobody is watching.

cancellation.py (illustrative pseudocode)
def stream_response(prompt, is_cancelled):
    for token in model.generate_stream(prompt):
        if is_cancelled():
            model.stop_generation()
            break
        yield token

Step 4 — Understand Backpressure

What are we doing? Recognizing that a client might not be able to consume tokens as fast as they're generated — a slow network connection, a slow renderer. Why? Without accounting for this, tokens can queue up unboundedly on the server, or get dropped.

Key Idea

Backpressure means the slower side of a pipeline sets the actual pace — a streaming system needs to either buffer safely or signal the producer to slow down.

Step 5 — Handle a Mid-Stream Error

What are we doing? Deciding what happens when generation fails after some tokens have already reached the client. Why? The client has already rendered a partial response — silently truncating it without explanation is a poor experience.

mid_stream_error.py (illustrative pseudocode)
def stream_response(prompt):
    try:
        for token in model.generate_stream(prompt):
            yield {"type": "token", "value": token}
    except ProviderError:
        yield {"type": "error", "message": "Generation was interrupted. Please try again."}

Step 6 — Note the Structured Output Caveat

A partial JSON object isn't valid JSON — streaming structured output needs a different strategy than streaming plain text, such as waiting for completion or streaming into a schema-aware incremental parser.

  • Assuming streaming reduces total generation time

    It changes when output becomes visible, not how fast the model generates it in total.

  • No cancellation handling

    The backend can keep generating and consuming resources for a response no one is watching.

  • No error handling for a failure partway through the stream

    The client may already show a partial response and needs a clear signal, not a silent cutoff.

  • Streaming structured output without special handling

    A partial JSON object isn't parseable.

  • Ignoring backpressure

    A slow client can cause unbounded buffering on the server if this isn't accounted for.

Challenges

Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.

Challenge 1: Add a heartbeat

Send a periodic no-op signal during long pauses so the client can detect a truly dead connection.

Challenge 2: Handle reconnection

Design how a client could resume a stream after a brief network drop.

Challenge 3: Stream a structured field incrementally

Attempt to stream one field of a JSON response as it completes, rather than waiting for the whole object.

Design Review

Before moving on, think through these questions the way a reviewer would.

  • What happens to server resources if 1,000 clients disconnect mid-stream at the same time?

  • How would you show the user that a stream was interrupted versus completed normally?

  • Where would caching interact with a streamed response?

Interview Questions

What does streaming actually improve, and what doesn't it improve?

It improves time to first token and perceived responsiveness — the user sees output start appearing quickly. It doesn't reduce the total time to generate the full response, since the model still produces tokens at the same underlying rate; streaming just exposes that process incrementally instead of waiting for it to finish.

  • Time to first token, not total time
  • Perceived vs. actual latency

How would you handle a client disconnecting in the middle of a stream?

I'd detect the disconnection and stop generation on the server side rather than letting it run to completion for no one — continuing to generate after the client is gone wastes provider resources and cost for output that will never be used.

  • Detect disconnection
  • Stop generation server-side
  • Avoid wasted cost

Why is streaming structured output harder than streaming plain text?

Because a partial JSON object isn't valid JSON — you can't parse or use it until it's complete. Streaming plain text to a UI works fine incrementally, but structured output generally needs to either wait for completion or use a schema-aware incremental parser designed for that purpose.

  • Partial JSON isn't parseable
  • Different strategy needed than plain text

Explain It in 30 Seconds

This project applies the streaming lesson to the full mechanics: sending each token to the client as it's generated, stopping generation when the client disconnects, and handling a failure partway through with a clear signal rather than a silent cutoff. Streaming improves time to first token and perceived responsiveness, not total generation time, and structured output needs different handling since a partial JSON object isn't valid.

On this page