LLM Gateway
Apply the gateway lessons to a working routing, retry, fallback, and cost-tracking layer between applications and model providers.
What You Will Build
An illustrative gateway that sits between application code and model providers, centralizing authentication, routing, retries, fallback, rate limiting, and cost tracking — so application code just makes a request and lets the gateway handle the rest. This project applies the LLM Gateway, AI Gateway Architecture, Rate Limiting, Cost Optimization, and Observability lessons.
Learning Objectives
Implement provider abstraction behind a single interface
Implement model routing based on task type and cost
Implement retry and fallback for a failed provider call
Add basic rate limiting and cost tracking
Prerequisites
Concepts Used
Architecture
Application
Makes One RequestLets the gateway handle the rest.
LLM Gateway (Auth, Routing, Rate Limit, Retry, Fallback, Logging)
Centralizes EverythingOne consistent interface behind every provider.
Provider A / Provider B
InterchangeableRouted to and swapped without changing application code.
Step 1 — Abstract the Providers
What are we doing? Presenting one consistent interface to application code, regardless of which provider actually handles a request. Why? Without this, provider-specific logic ends up duplicated across every calling application.
class Provider:
def call(self, request): ...
class ProviderA(Provider):
def call(self, request):
return provider_a_sdk.generate(request.to_provider_a_format())
class ProviderB(Provider):
def call(self, request):
return provider_b_sdk.complete(request.to_provider_b_format())Step 2 — Route by Task
What are we doing? Choosing which model handles a request, based on task type rather than a hardcoded choice in application code. Why? Simple tasks can use a cheaper, faster model; complex reasoning may need a stronger one — this is a cost and latency lever, not just a technical choice.
def select_provider(task_type):
if task_type == "classification":
return ProviderA() # fast, cheap
if task_type == "reasoning":
return ProviderB() # stronger, slower
return ProviderA()Step 3 — Add Retry and Fallback
What are we doing? Retrying a failed call, and falling back to an alternate provider if the primary one is unavailable. Why? Without a fallback, a single provider outage becomes a full outage for every application behind the gateway.
def handle_request(request):
provider = select_provider(request.task_type)
try:
return call_with_retry(provider, request, max_retries=2, timeout=request.timeout)
except ProviderUnavailable:
fallback = get_fallback(provider)
return call_with_retry(fallback, request, max_retries=1, timeout=request.timeout)
except AllProvidersFailed:
return degraded_response(request)Step 4 — Add Rate Limiting and Cost Tracking
What are we doing? Bounding how much any one caller can use, and recording enough to understand spend. Why? Without limits, one misbehaving caller can exhaust shared provider quota; without tracking, cost problems are only discovered when the bill arrives.
Request
Incoming CallFrom an application behind the gateway.
Rate Limiter
Bounds UsageStops one caller from exhausting shared quota.
Provider Call
Actual RequestSent to the routed provider.
Log Tokens + Cost + Latency
Spend VisibilityCaught early instead of at the bill.
Step 5 — Add Observability
What are we doing? Logging enough per request to debug and track cost, without capturing unnecessary sensitive content. Why? The gateway sees every request — it's the natural place to centralize this visibility, but also a common place for sensitive data to leak if logged indiscriminately.
No fallback when the primary provider fails
A single provider outage becomes a full outage for every application behind the gateway.
No per-caller rate limiting
One misbehaving caller can exhaust shared provider quota for everyone.
Logging full prompts and responses without redaction
Gateway logs are a common place for sensitive content to leak.
Hardcoding model choice in application code
This defeats the purpose of routing — the decision should live in the gateway.
Making the gateway itself a single point of failure
It needs its own reliability story, since every request now depends on it.
Challenges
Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.
Challenge 1: Add a third provider
Add a second fallback tier so the gateway tries a third provider if both primary and first fallback fail.
Challenge 2: Add per-tenant rate limits
Extend rate limiting so each tenant has its own independent quota.
Challenge 3: Add cost budgets
Add a per-tenant monthly cost budget that routes to a cheaper model once exceeded.
Design Review
Before moving on, think through these questions the way a reviewer would.
What happens to every application behind this gateway if the gateway itself goes down?
How would you decide when a request should retry versus immediately fall back?
Where would you add caching to reduce load on the providers?
How would you detect a sudden cost spike before the monthly bill arrives?
Interview Questions
How would you design an LLM Gateway, and when would you actually introduce one?
I'd introduce a gateway once multiple applications or models create real duplication in authentication, retries, and cost tracking. It should abstract provider differences behind one interface, route by task type, retry and fall back to an alternate provider on failure, apply rate limits, and log enough for cost tracking and debugging without capturing sensitive content indiscriminately. For a single small app calling one model, I'd skip it — it's an extra hop and operational component that isn't earning its cost yet.
- Introduce once duplication is real
- Provider abstraction, routing, retry/fallback
- Careful, non-indiscriminate logging
How would you handle a model provider outage?
I'd have the gateway detect the failure — timeout or explicit error — and automatically route to a configured fallback provider or model, logging which path actually served the request. If all providers fail, I'd return a degraded response rather than letting the failure propagate as a hard error to every calling application.
- Automatic fallback
- Graceful degradation as last resort
- Visibility into which path served the request
Why is model routing also a cost lever, not just a performance one?
Because routing simple, low-stakes requests to a smaller, cheaper model directly reduces spend, while reserving a stronger, more expensive model for tasks that actually need it. It's one of the biggest cost levers available, since it changes cost per request rather than just optimizing infrastructure.
- Model choice drives cost directly
- Bigger lever than infrastructure tuning
Explain It in 30 Seconds
This project applies the gateway lessons to a working layer between applications and model providers: provider abstraction behind one interface, routing by task type, retry and fallback on failure, rate limiting, and cost/latency logging. It's worth the added hop and operational overhead once multiple applications or models create real duplication — and it needs its own reliability story, since every request now depends on it.