Rate Limiting
Rate limiting restricts how many requests a client can make in a given time window to protect a system from overload.
Why Rate Limiting Matters for AI Applications
AI applications have two rate limits to think about at once: the limits your own system sets for its users, and the limits the model provider imposes on you. Exceeding either one causes requests to fail — but the provider's limit is shared across your entire application, so one misbehaving user or feature can affect everyone if there's no limiting in front of it.
Incoming Requests
From All UsersTraffic arriving from every client at once.
Limiter
Enforces LimitsChecks each request against a rate ceiling.
AI Service
Protected ResourceThe shared, provider-limited model call.
Rejected / Delayed
Over LimitFails fast rather than overwhelming the provider.
Common Strategies
- Fixed window
- Allow up to N requests per fixed time period (like per minute) — simple, but can allow a burst right at the boundary between two windows.
- Sliding window
- Smooths out that boundary issue by considering a rolling window rather than fixed clock-aligned periods.
- Token bucket
- A bucket that refills at a steady rate and is drained by each request — naturally allows some burst capacity while enforcing a long-run average rate.
- Concurrency limits
- Limiting how many requests can be in flight at once, rather than (or alongside) how many happen per time period — often more directly relevant for expensive, slow LLM calls.
Layers That Need Limiting
- Per-user limits — preventing a single user from consuming disproportionate resources or cost.
- Per-tenant limits — in a multi-tenant system, preventing one tenant from degrading service for others.
- Global limits toward the provider — respecting the provider's own rate limit for your application as a whole, since exceeding it can fail requests for every user at once.
- Retry behavior — a client retrying immediately and aggressively after being rate-limited can make the problem worse; retries should back off, not compound the load.
Warning
A retry storm — many clients retrying a rate-limited or failing request at the same time — can turn a temporary slowdown into a full outage.
Common Mistakes
Only limiting at one layer
Per-user, per-tenant, and provider-level limits solve different problems — missing one leaves a real gap.
Retrying immediately and aggressively after a rate limit error
This is exactly how a retry storm starts — retries should use backoff, not repeat instantly.
Ignoring the provider's own rate limits until they're hit in production
Provider limits are a hard external constraint — they should be designed around proactively, not discovered through failures.
Using time-based rate limiting alone for expensive LLM calls
A concurrency limit is often more directly relevant, since a small number of slow, expensive requests can saturate capacity faster than a larger number of fast ones.
Interview Question
How would you design rate limiting for an AI application that calls an external model provider?
I'd limit at several layers, since they solve different problems: per-user limits to stop one user from consuming disproportionate cost, per-tenant limits in a multi-tenant system so one tenant can't degrade service for others, and a global limit that respects the provider's own rate limit, since exceeding that can fail requests for the whole application at once. For LLM calls specifically, I'd also consider a concurrency limit rather than only a time-based one, since a small number of slow, expensive calls can saturate capacity faster than many fast ones. Retries need backoff, not immediate retry, because aggressive retrying after a rate limit error is exactly how a retry storm turns a temporary slowdown into a full outage.
What an interviewer may ask next
- Why might you need rate limiting at multiple layers rather than just one?
- What is a retry storm, and how do you prevent one?
- Why is a concurrency limit sometimes more relevant than a time-based rate limit for LLM calls?
Explain It in 30 Seconds
Rate limiting for AI applications has to account for both your own system's limits and the model provider's shared limits — exceeding either fails requests, but the provider's limit affects your whole application at once. It's usually applied at several layers: per-user, per-tenant, and globally toward the provider, often combined with a concurrency limit since slow, expensive LLM calls can saturate capacity differently than fast ones. Retries need backoff, since aggressive immediate retries after a rate limit error are exactly how a retry storm starts.