AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Advanced45–60 min

Multi-Tenant AI System

Apply the multi-tenant architecture lessons to tenant isolation, quotas, and cost allocation on top of a shared AI system.

What You Will Build

An illustrative multi-tenant layer on top of a shared AI system: resolving a request's tenant, enforcing data isolation at retrieval time, applying per-tenant rate limits and quotas, and tracking cost per tenant. This project applies Multi-Tenant AI Architecture, AI Security, LLM Gateway, Rate Limiting, and Cost Optimization.

Learning Objectives

  • Enforce tenant isolation at the retrieval layer, not the model

  • Apply per-tenant rate limits and usage tracking

  • Understand the tradeoff between logical and physical isolation

  • Understand how cost allocation works per tenant

Prerequisites

Concepts Used

Multi-Tenant AI Architecture
AI Security
LLM Gateway
Rate Limiting
Cost Optimization
Metadata

Architecture

Request

Incoming Call

Before tenant identity is known.

passes through

Authentication

Confirms Identity

Runs before tenant resolution.

then

Resolve Tenant

Every Decision Depends On This

Determines data access and quota.

scopes

Tenant-Scoped Retrieval

Isolation Enforced Here

Filtered by tenant, not left to the model.

feeds into

AI Application (LLM / RAG / Tools)

Shared Infrastructure

Same system serving every tenant.

produces

Response

Tenant-Scoped Result

Only reflects that tenant's own data.

Multi-tenant AI request

Step 1 — Resolve the Tenant

What are we doing? Determining which tenant a request belongs to, right after authentication. Why? Every downstream decision — what data this request can access, what quota applies — depends on knowing the tenant first.

resolve_tenant.py (illustrative pseudocode)
def handle_request(request):
    user = authenticate(request)
    tenant_id = resolve_tenant(user)
    return process(request, tenant_id)

Step 2 — Enforce Isolation at Retrieval

What are we doing? Filtering every retrieval query by the resolved tenant. Why? A model has no inherent concept of tenant boundaries — isolation has to be enforced by the retrieval layer, not left to the model's judgment.

tenant_retrieval.py (illustrative pseudocode)
def retrieve(query, tenant_id):
    return vector_store.search(query, filter={"tenant": tenant_id})

Important

A single unfiltered retrieval path can leak another tenant's content, even if every other path in the system is correct.

Step 3 — Choose an Isolation Strategy

Logical isolation

One shared index

Tenant ID on every record

Cheaper to operate

Physical isolation

Separate index per tenant

Isolation enforced structurally

More operational overhead

Most systems start with logical isolation and move specific tenants to physical isolation only when a contractual or regulatory requirement demands it.

Step 4 — Add Per-Tenant Quotas

What are we doing? Limiting usage per tenant, not just globally. Why? Without this, one tenant's heavy usage can degrade latency or exhaust shared quota for every other tenant — a noisy-neighbor problem.

Step 5 — Track Cost Per Tenant

What are we doing? Recording token usage and cost attributed to each tenant. Why? Without this, cost can't be allocated, billed, or investigated per tenant — it just shows up as one combined number.

  • Relying on the model to keep tenants separate

    A model has no inherent concept of tenant boundaries — isolation must be enforced by the application.

  • Forgetting tenant scoping on a shared cache

    A cache key without a tenant identifier can serve one tenant's cached response to another.

  • No per-tenant rate limiting

    One tenant can degrade service or dominate cost for everyone else on shared infrastructure.

  • Inconsistent metadata tagging at ingestion

    If tenant metadata isn't reliably attached when data is ingested, filtering at query time becomes unreliable.

  • Assuming logical isolation is sufficient for every tenant

    Some tenants may require physical isolation for regulatory or contractual reasons.

Challenges

Extend the project yourself. No automated grading — use these to practice reasoning about the architecture.

Challenge 1: Add tenant-level configuration

Let each tenant have a different system prompt or allowed tool set.

Challenge 2: Add an audit log

Log every data access with tenant, user, and timestamp so access can be verified after the fact.

Challenge 3: Migrate one tenant to physical isolation

Design what changes to migrate a single high-requirement tenant from the shared index to a dedicated one.

Design Review

Before moving on, think through these questions the way a reviewer would.

  • How would you verify, after the fact, that no cross-tenant data leak occurred?

  • What would justify moving a specific tenant from logical to physical isolation?

  • How would you prevent one tenant's usage spike from degrading service for others?

  • Where does tenant identity need to be checked more than once in this flow?

Interview Questions

How would you design tenant isolation in a shared RAG or AI system?

I'd enforce isolation at the retrieval layer, not rely on the model — every document and cache entry gets a tenant identifier at write time, and every read filters by the requesting tenant. I'd start with logical isolation, one shared index with tenant metadata on every record, and move to physical isolation only where a specific tenant's contractual or regulatory needs require it. I'd also add per-tenant rate limiting and cost tracking so no tenant can degrade service or cost for others.

  • Enforced at retrieval, not the model
  • Logical isolation by default, physical when required
  • Per-tenant limits and cost tracking

What could go wrong with a shared cache in a multi-tenant system?

If the cache key doesn't include a tenant identifier, a response computed for one tenant could be served to another — a direct data leak. Every cache key needs to be scoped by tenant, the same way retrieval queries are.

  • Cache keys need tenant scoping
  • Same root issue as retrieval isolation

How would you allocate AI cost across tenants?

I'd track token usage and cost per request, tagged with the resolved tenant, at the gateway or application layer — the same place routing and rate limiting already happen. That gives a per-tenant cost total that can be billed, budgeted, or investigated, instead of one combined number that hides which tenant is driving cost.

  • Tag usage with tenant at the point of the model call
  • Enables billing and cost investigation

Explain It in 30 Seconds

This project applies multi-tenant architecture to a shared AI system: resolving the tenant right after authentication, enforcing isolation at the retrieval layer since a model has no inherent concept of tenant boundaries, and applying per-tenant rate limits and cost tracking. Most systems start with logical isolation — one shared index with tenant metadata — and move to physical isolation only where a specific tenant genuinely requires it.

On this page