AI Gateway Architecture
AI gateway architecture describes how requests are routed, authenticated, and load balanced across model providers.
Prerequisites
What a Gateway Centralizes
As soon as more than one application or more than one model provider is involved, calling providers directly from application code means duplicating the same concerns everywhere: authentication, retries, rate limiting, logging. An AI Gateway sits between applications and model providers and centralizes those concerns in one place, so application code just makes a request and lets the gateway handle the rest.
Application
Makes RequestJust asks the gateway, without knowing which provider.
AI Gateway
Centralizes ConcernsHandles auth, retries, rate limiting, and logging.
Provider A (fast model)
Low-Latency RouteUsed for quick, simple requests.
Provider B (reasoning model)
Complex RouteUsed when a task needs deeper reasoning.
Provider C (embeddings)
Embeddings RouteA separate provider dedicated to embeddings.
Key Idea
A gateway turns 'every application knows how to call every provider' into 'every application knows how to call the gateway.'
Model Routing
Routing is one of the gateway's most valuable responsibilities: choosing which model actually handles a given request, based on the task rather than a hardcoded choice in application code.
- Task type — simple classification or extraction can often use a smaller, cheaper, faster model; complex reasoning may need a stronger one.
- Latency requirements — an interactive chat response has a tighter latency budget than a background summarization job.
- Cost — routing lower-stakes requests to cheaper models is a direct cost control, not just a technical choice.
- Availability — if a preferred provider is degraded or unavailable, the gateway can route to a fallback rather than failing the request outright.
Code Example
Illustrative pseudocode — a real gateway typically has far more configuration, but the routing and fallback shape looks like this:
def route_request(request):
model = select_model(request.task_type, request.latency_budget)
try:
return call_with_retry(model, request, max_retries=2, timeout=request.timeout)
except ProviderUnavailable:
fallback_model = get_fallback(model)
return call_with_retry(fallback_model, request, max_retries=1, timeout=request.timeout)
except AllProvidersFailed:
return degraded_response(request)
def select_model(task_type, latency_budget):
if task_type == "classification":
return "fast-model"
if task_type == "reasoning" and latency_budget > SLOW_THRESHOLD:
return "reasoning-model"
return "default-model"Tradeoffs
One less network hop
Simple for a single app, single model
Logic duplicated across apps as usage grows
Centralized auth, routing, retries, cost tracking
One place to add a new provider or model
Added latency hop and an operational component to run
Warning
A gateway isn't automatically the right choice. For a single application calling a single model, it can be pure overhead — introduce it when multiple applications, multiple models, or genuine cross-cutting concerns justify the centralization.
Common Mistakes
Introducing a gateway before there's a real need
A single app calling a single provider doesn't need centralization yet — it just adds a network hop and an operational component.
No fallback when the primary provider fails
Without a fallback path, a single provider outage becomes a full outage for every application behind the gateway.
No per-request or per-tenant rate limiting
Without limits, one misbehaving caller can exhaust shared provider quota for everyone behind the gateway.
Logging full prompts and responses without redaction
Gateway logs are a natural place for sensitive content to leak if request and response bodies are logged indiscriminately.
Making the gateway itself a single point of failure
The gateway needs its own reliability story — redundancy, health checks — since every request now depends on it.
Hardcoding model choice in application code
This defeats the purpose of routing — the model decision should live in the gateway, not scattered across every calling application.
Interview Question
How would you design an AI Gateway, and when would you actually introduce one?
I'd introduce an AI Gateway once multiple applications or multiple models create real duplication — every app otherwise has to implement its own authentication, retries, rate limiting, and cost tracking against each provider. The gateway centralizes those concerns and adds routing: choosing which model handles a request based on task type, latency budget, and cost, with a fallback path if the preferred provider is degraded or unavailable. For a single small application calling one model directly, I'd avoid it — it's an extra network hop and an operational component that isn't earning its cost yet. Wherever I do introduce it, the gateway needs its own reliability story, since every request now depends on it, and its logs need to avoid indiscriminately capturing sensitive prompt or response content.
What an interviewer may ask next
- What would you route on besides just task type?
- What happens if the gateway itself becomes unavailable?
- Why might logging full prompts and responses at the gateway be risky?
- When would you NOT introduce an AI Gateway?
Explain It in 30 Seconds
An AI Gateway sits between applications and model providers, centralizing authentication, routing, retries, fallback, rate limiting, and cost tracking so application code doesn't duplicate that logic per provider. Routing chooses which model handles a request based on task type, latency budget, cost, and availability, with a fallback if the preferred provider fails. It's worth the added network hop and operational overhead once multiple applications or models create real duplication — not for a single app calling a single model.