Streaming AI Application
Apply the Streaming lesson to the full mechanics: time to first token, cancellation, backpressure, and mid-stream errors.
What You Will Build
An illustrative streaming pipeline that sends model output incrementally to a client, handles a client disconnecting mid-stream, and handles a failure partway through generation. No real LLM provider is required — the streaming mechanics are illustrated directly.
Learning Objectives
Understand why streaming improves perceived latency without reducing total generation time
Implement cancellation when a client disconnects mid-stream
Handle an error that occurs after some tokens have already been sent
Understand backpressure and why a client can't always keep up
Prerequisites
Concepts Used
Architecture
Request
Starts GenerationKicks off the model call.
LLM Generates Token
Same Underlying RateStreaming doesn't speed this up.
Send to Client
ImmediatelyAs soon as each token exists.
Client Renders
Time to First TokenWhat actually improves perceived speed.
Repeat Until Done or Cancelled
Loop ExitStops on completion or client disconnect.
Step 1 — Compare Normal vs. Streaming
What are we doing? Contrasting a request that waits for the full response with one that streams incrementally. Why? This makes the actual improvement — time to first token, not total time — concrete before adding the more complex handling.
Request
Wait
Complete Response
Request
Token 1
Token 2
Token 3
…
Complete
Step 2 — Implement a Basic Stream
What are we doing? Sending each generated token to the client as soon as it exists. How it works: illustrated here as a generator that yields tokens one at a time, standing in for a real model's token-by-token output.
def stream_response(prompt):
for token in model.generate_stream(prompt):
yield token # sent to the client immediately
# Client side: render each token as it arrives
for token in stream_response(prompt):
append_to_ui(token)Step 3 — Handle Cancellation
What are we doing? Stopping generation when the client disconnects or the user navigates away. Why? Without this, the backend keeps generating and consuming provider resources for a response nobody is watching.
def stream_response(prompt, is_cancelled):
for token in model.generate_stream(prompt):
if is_cancelled():
model.stop_generation()
break
yield tokenStep 4 — Understand Backpressure
What are we doing? Recognizing that a client might not be able to consume tokens as fast as they're generated — a slow network connection, a slow renderer. Why? Without accounting for this, tokens can queue up unboundedly on the server, or get dropped.
Key Idea
Backpressure means the slower side of a pipeline sets the actual pace — a streaming system needs to either buffer safely or signal the producer to slow down.
Step 5 — Handle a Mid-Stream Error
What are we doing? Deciding what happens when generation fails after some tokens have already reached the client. Why? The client has already rendered a partial response — silently truncating it without explanation is a poor experience.
def stream_response(prompt):
try:
for token in model.generate_stream(prompt):
yield {"type": "token", "value": token}
except ProviderError:
yield {"type": "error", "message": "Generation was interrupted. Please try again."}Step 6 — Note the Structured Output Caveat
A partial JSON object isn't valid JSON — streaming structured output needs a different strategy than streaming plain text, such as waiting for completion or streaming into a schema-aware incremental parser.
Assuming streaming reduces total generation time
It changes when output becomes visible, not how fast the model generates it in total.
No cancellation handling
The backend can keep generating and consuming resources for a response no one is watching.
No error handling for a failure partway through the stream
The client may already show a partial response and needs a clear signal, not a silent cutoff.
Streaming structured output without special handling
A partial JSON object isn't parseable.
Ignoring backpressure
A slow client can cause unbounded buffering on the server if this isn't accounted for.
Challenges
Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.
Challenge 1: Add a heartbeat
Send a periodic no-op signal during long pauses so the client can detect a truly dead connection.
Challenge 2: Handle reconnection
Design how a client could resume a stream after a brief network drop.
Challenge 3: Stream a structured field incrementally
Attempt to stream one field of a JSON response as it completes, rather than waiting for the whole object.
Design Review
Before moving on, think through these questions the way a reviewer would.
What happens to server resources if 1,000 clients disconnect mid-stream at the same time?
How would you show the user that a stream was interrupted versus completed normally?
Where would caching interact with a streamed response?
Interview Questions
What does streaming actually improve, and what doesn't it improve?
It improves time to first token and perceived responsiveness — the user sees output start appearing quickly. It doesn't reduce the total time to generate the full response, since the model still produces tokens at the same underlying rate; streaming just exposes that process incrementally instead of waiting for it to finish.
- Time to first token, not total time
- Perceived vs. actual latency
How would you handle a client disconnecting in the middle of a stream?
I'd detect the disconnection and stop generation on the server side rather than letting it run to completion for no one — continuing to generate after the client is gone wastes provider resources and cost for output that will never be used.
- Detect disconnection
- Stop generation server-side
- Avoid wasted cost
Why is streaming structured output harder than streaming plain text?
Because a partial JSON object isn't valid JSON — you can't parse or use it until it's complete. Streaming plain text to a UI works fine incrementally, but structured output generally needs to either wait for completion or use a schema-aware incremental parser designed for that purpose.
- Partial JSON isn't parseable
- Different strategy needed than plain text
Explain It in 30 Seconds
This project applies the streaming lesson to the full mechanics: sending each token to the client as it's generated, stopping generation when the client disconnects, and handling a failure partway through with a clear signal rather than a silent cutoff. Streaming improves time to first token and perceived responsiveness, not total generation time, and structured output needs different handling since a partial JSON object isn't valid.