LLM Gateway
An LLM gateway is a shared layer that routes requests to model providers while handling auth, logging, and fallback.
Prerequisites
Design vs. Implementation
AI Gateway Architecture explains why and when a gateway pattern makes sense at a system-design level. This lesson is about actually building one: the concrete responsibilities and code paths an LLM gateway needs to implement, whether you build it yourself or adopt an existing gateway product.
Application
Single CallMakes one request without knowing the provider details.
LLM Gateway
Shared LayerOne place implementing gateway responsibilities.
Auth + Config
Credentials ResolvedHandles authentication and provider settings.
Routing
Picks a ProviderDecides which model provider handles this request.
Provider A / Provider B
Actual Model CallWhichever provider was chosen, with fallback available.
Key Idea
"LLM gateway" and "AI gateway" refer to the same component — this lesson uses "LLM gateway" since it's the more common term for the concrete, buildable layer described here.
Core Responsibilities
- Provider abstraction
- Presenting one consistent interface to application code, regardless of which underlying provider actually handles a given request.
- Authentication
- Holding provider credentials centrally, so individual applications never need their own copies of provider API keys.
- Model routing
- Choosing which model handles a request, based on task type, cost, or availability — covered in more system-design depth in AI Gateway Architecture.
- Retries and fallback
- Automatically retrying a failed call, and falling back to an alternate model or provider if the primary one is unavailable.
- Logging and cost tracking
- Recording enough about each request — model, latency, token counts — to understand usage and cost, without capturing unnecessary sensitive content.
- Configuration
- Centralizing settings like default models, timeouts, and rate limits so they can be changed in one place rather than across many applications.
Code Example
Illustrative pseudocode showing the shape of a gateway request handler:
def handle_request(request, config):
provider = config.route(request.task_type)
start = time.time()
try:
response = provider.call(request, timeout=config.timeout)
log_request(request, provider, latency=time.time() - start, status="ok")
return response
except ProviderError:
fallback = config.fallback_for(provider)
response = fallback.call(request, timeout=config.timeout)
log_request(request, fallback, latency=time.time() - start, status="fallback")
return responseCommon Mistakes
Letting application code hold provider credentials directly
Centralizing authentication in the gateway means credentials live in one place, not scattered across every calling application.
No fallback path when the primary provider fails
Without one, a single provider outage becomes a full outage for every application behind the gateway.
Logging full prompt and response content indiscriminately
Gateway logs are a common place for sensitive data to leak — log what's needed for debugging and cost tracking, not everything by default.
Hardcoding provider-specific logic in application code instead of the gateway
This defeats the point of provider abstraction — the gateway should be the only place that knows provider-specific details.
No timeout on the underlying provider call
Without one, a slow provider can hang the request far longer than acceptable for the application calling it.
Interview Question
What are the core responsibilities of an LLM gateway, and how do you implement retries and fallback?
An LLM gateway needs to abstract away provider differences behind one consistent interface, hold authentication centrally so credentials aren't scattered across applications, route requests to the right model, retry and fall back to an alternate provider on failure, log enough to track usage and cost without capturing unnecessary sensitive content, and centralize configuration like timeouts and defaults. For retries and fallback specifically, I'd wrap the provider call with a timeout, retry a bounded number of times on transient failures, and if the primary provider still fails, route to a configured fallback provider or model — logging which path actually served the request so it's visible in observability.
What an interviewer may ask next
- Why should provider credentials live in the gateway rather than in each calling application?
- What should the gateway log, and what should it avoid logging?
- How would you decide when to retry a failed call versus immediately falling back to another provider?
Explain It in 30 Seconds
An LLM gateway is the concrete implementation of provider abstraction, authentication, model routing, retries and fallback, logging, and configuration — one shared layer application code calls instead of talking to providers directly. It centralizes credentials so they're not scattered across applications, and gives every application consistent retry and fallback behavior without duplicating that logic. Logging needs to capture enough for cost tracking and debugging without indiscriminately recording sensitive prompt content.