Multi-Tenant AI System
Apply the multi-tenant architecture lessons to tenant isolation, quotas, and cost allocation on top of a shared AI system.
What You Will Build
An illustrative multi-tenant layer on top of a shared AI system: resolving a request's tenant, enforcing data isolation at retrieval time, applying per-tenant rate limits and quotas, and tracking cost per tenant. This project applies Multi-Tenant AI Architecture, AI Security, LLM Gateway, Rate Limiting, and Cost Optimization.
Learning Objectives
Enforce tenant isolation at the retrieval layer, not the model
Apply per-tenant rate limits and usage tracking
Understand the tradeoff between logical and physical isolation
Understand how cost allocation works per tenant
Prerequisites
Concepts Used
Architecture
Request
Incoming CallBefore tenant identity is known.
Authentication
Confirms IdentityRuns before tenant resolution.
Resolve Tenant
Every Decision Depends On ThisDetermines data access and quota.
Tenant-Scoped Retrieval
Isolation Enforced HereFiltered by tenant, not left to the model.
AI Application (LLM / RAG / Tools)
Shared InfrastructureSame system serving every tenant.
Response
Tenant-Scoped ResultOnly reflects that tenant's own data.
Step 1 — Resolve the Tenant
What are we doing? Determining which tenant a request belongs to, right after authentication. Why? Every downstream decision — what data this request can access, what quota applies — depends on knowing the tenant first.
def handle_request(request):
user = authenticate(request)
tenant_id = resolve_tenant(user)
return process(request, tenant_id)Step 2 — Enforce Isolation at Retrieval
What are we doing? Filtering every retrieval query by the resolved tenant. Why? A model has no inherent concept of tenant boundaries — isolation has to be enforced by the retrieval layer, not left to the model's judgment.
def retrieve(query, tenant_id):
return vector_store.search(query, filter={"tenant": tenant_id})Important
A single unfiltered retrieval path can leak another tenant's content, even if every other path in the system is correct.
Step 3 — Choose an Isolation Strategy
One shared index
Tenant ID on every record
Cheaper to operate
Separate index per tenant
Isolation enforced structurally
More operational overhead
Most systems start with logical isolation and move specific tenants to physical isolation only when a contractual or regulatory requirement demands it.
Step 4 — Add Per-Tenant Quotas
What are we doing? Limiting usage per tenant, not just globally. Why? Without this, one tenant's heavy usage can degrade latency or exhaust shared quota for every other tenant — a noisy-neighbor problem.
Step 5 — Track Cost Per Tenant
What are we doing? Recording token usage and cost attributed to each tenant. Why? Without this, cost can't be allocated, billed, or investigated per tenant — it just shows up as one combined number.
Relying on the model to keep tenants separate
A model has no inherent concept of tenant boundaries — isolation must be enforced by the application.
Forgetting tenant scoping on a shared cache
A cache key without a tenant identifier can serve one tenant's cached response to another.
No per-tenant rate limiting
One tenant can degrade service or dominate cost for everyone else on shared infrastructure.
Inconsistent metadata tagging at ingestion
If tenant metadata isn't reliably attached when data is ingested, filtering at query time becomes unreliable.
Assuming logical isolation is sufficient for every tenant
Some tenants may require physical isolation for regulatory or contractual reasons.
Challenges
Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.
Challenge 1: Add tenant-level configuration
Let each tenant have a different system prompt or allowed tool set.
Challenge 2: Add an audit log
Log every data access with tenant, user, and timestamp so access can be verified after the fact.
Challenge 3: Migrate one tenant to physical isolation
Design what changes to migrate a single high-requirement tenant from the shared index to a dedicated one.
Design Review
Before moving on, think through these questions the way a reviewer would.
How would you verify, after the fact, that no cross-tenant data leak occurred?
What would justify moving a specific tenant from logical to physical isolation?
How would you prevent one tenant's usage spike from degrading service for others?
Where does tenant identity need to be checked more than once in this flow?
Interview Questions
How would you design tenant isolation in a shared RAG or AI system?
I'd enforce isolation at the retrieval layer, not rely on the model — every document and cache entry gets a tenant identifier at write time, and every read filters by the requesting tenant. I'd start with logical isolation, one shared index with tenant metadata on every record, and move to physical isolation only where a specific tenant's contractual or regulatory needs require it. I'd also add per-tenant rate limiting and cost tracking so no tenant can degrade service or cost for others.
- Enforced at retrieval, not the model
- Logical isolation by default, physical when required
- Per-tenant limits and cost tracking
What could go wrong with a shared cache in a multi-tenant system?
If the cache key doesn't include a tenant identifier, a response computed for one tenant could be served to another — a direct data leak. Every cache key needs to be scoped by tenant, the same way retrieval queries are.
- Cache keys need tenant scoping
- Same root issue as retrieval isolation
How would you allocate AI cost across tenants?
I'd track token usage and cost per request, tagged with the resolved tenant, at the gateway or application layer — the same place routing and rate limiting already happen. That gives a per-tenant cost total that can be billed, budgeted, or investigated, instead of one combined number that hides which tenant is driving cost.
- Tag usage with tenant at the point of the model call
- Enables billing and cost investigation
Explain It in 30 Seconds
This project applies multi-tenant architecture to a shared AI system: resolving the tenant right after authentication, enforcing isolation at the retrieval layer since a model has no inherent concept of tenant boundaries, and applying per-tenant rate limits and cost tracking. Most systems start with logical isolation — one shared index with tenant metadata — and move to physical isolation only where a specific tenant genuinely requires it.