Streaming
Streaming returns a model's output incrementally as it's generated instead of waiting for the full response.
Why Streaming Exists
A decoder generates a response one token at a time. Without streaming, an application waits for every one of those tokens before showing anything to the user — for a long response, that can mean many seconds of a blank screen. Streaming sends each token to the client as soon as it's generated, so the user sees the response appear progressively instead of waiting for the whole thing.
LLM Generates Token
One at a TimeThe decoder produces a single token.
Token Sent Immediately
No WaitingSent right away, not batched with the rest.
Client Renders Token
Progressive DisplayThe user sees the response appear as it forms.
Repeat Until Done
ContinuesRuns again for every remaining token.
Key Idea
Streaming doesn't make the model faster overall — it changes when the user starts seeing output, which is what "time to first token" measures.
Request/Response vs. Streaming
Send request
Wait for full generation
Receive complete response
Simple to implement
Send request
Receive tokens incrementally
Render as they arrive
Better perceived latency
Request/response is simpler and fine for short outputs or backend-to-backend calls where nothing is rendered live. Streaming matters most for user-facing, longer responses, where perceived responsiveness — how quickly something starts happening — matters as much as total completion time.
Engineering Considerations
- Cancellation — a user might navigate away or stop generation mid-stream; the application needs a clean way to end the connection and stop consuming tokens it no longer needs.
- Errors mid-stream — a failure partway through generation needs a defined behavior, since the client has already rendered a partial response by that point.
- Buffering on the client — depending on the UI, tokens might need light buffering to render smoothly rather than flickering character by character.
- Structured output and streaming don't mix cleanly — a partial JSON object isn't valid JSON, so streaming structured output requires special handling.
Common Mistakes
Assuming streaming reduces total generation time
Streaming changes when output becomes visible, not how fast the model actually generates it in total.
No handling for a client disconnecting mid-stream
Without cancellation handling, the backend can keep generating and consuming provider resources for a response no one is watching anymore.
Streaming structured output without special handling
A partial JSON object isn't parseable — streaming structured output needs a strategy for handling incomplete data, not just streaming raw text.
No error handling for a failure partway through the stream
The client may already show a partial response — an application needs a defined way to signal that generation didn't complete successfully.
Using streaming for backend-to-backend calls with no UI
If nothing is rendering the output live, request/response is usually simpler and just as effective.
Interview Question
What is streaming, and what does it actually improve compared to a standard request/response call?
Streaming sends each token to the client as soon as it's generated, instead of waiting for the entire response — since a decoder produces output one token at a time anyway, streaming just exposes that incremental process to the user. It doesn't make the model generate faster overall; what it improves is time to first token and perceived responsiveness, which matters most for user-facing, longer responses. It does add real engineering considerations: handling a client disconnecting mid-stream, handling a failure partway through generation when a partial response is already showing, and the fact that structured output doesn't stream cleanly, since a partial JSON object isn't valid JSON.
What an interviewer may ask next
- Does streaming reduce the total time it takes a model to finish generating a response?
- What should happen if a client disconnects in the middle of a stream?
- Why is streaming structured output more complicated than streaming plain text?
Explain It in 30 Seconds
Streaming sends each token to the client as soon as it's generated, instead of waiting for the full response, since a decoder already produces output one token at a time. It doesn't reduce total generation time — it improves time to first token and perceived responsiveness for user-facing output. It adds real engineering work too: handling cancellation when a client disconnects, handling errors mid-stream, and the fact that structured output like JSON doesn't stream cleanly since a partial object isn't valid.