API Design for AI
AI APIs need to account for streaming, long-running requests, and partial or failed generations that a typical REST API does not.
Prerequisites
Overview
A typical REST endpoint returns one response for one request. An AI endpoint often needs to stream a partial response, support cancellation mid-generation, and report a graceful partial result if a model call fails halfway through.
Where It Fits
Request
Stream Starts
Partial Tokens
Complete or Cancelled/Failed
Key Points
- Streaming responses
- Server-sent events or chunked responses let a client render output before generation finishes.
- Cancellation
- A client should be able to signal cancellation, and the server should actually stop the underlying provider call, not just stop reading it.
- Partial failure handling
- If a stream fails midway, the API should return what was generated so far along with a clear error, not silently drop it.
Interview Question
How is designing an API for an LLM endpoint different from a typical CRUD REST API?
A CRUD endpoint returns one response for one request; an LLM endpoint usually streams a partial response over time, needs to support real cancellation of the underlying provider call, and needs a contract for partial failure — returning what generated so far plus a clear error, rather than just a generic 500.
Explain It in 30 Seconds
AI APIs differ from typical REST APIs mainly in three ways: they stream partial output over time, they need real cancellation of an in-flight generation, and they need to handle partial failure gracefully rather than treating a mid-stream error like any other request failure.
Real-World Stack
Technologies commonly used to implement this in production.