Kubernetes for AI
Kubernetes runs and scales the containers behind an AI service — the API layer, gateway, and any self-hosted models.
Overview
Kubernetes doesn’t know anything about models specifically — it schedules and scales containers. What makes it relevant to AI is that the API layer, an AI gateway, and any self-hosted model server are all just containers that need exactly that.
Where It Fits
Ingress / Load Balancer
API Pods
AI Gateway Pods
Self-Hosted Model Pods (GPU)
Key Points
- GPU scheduling
- Kubernetes needs GPU-aware node pools and resource requests to schedule model-serving pods onto the right hardware.
- Autoscaling nuance
- Standard CPU-based autoscaling doesn’t capture GPU or queue-depth load well — AI workloads often need custom scaling metrics.
- Stateless API layer
- The API and gateway layers scale easily since they’re typically stateless; the model-serving layer is the more constrained resource.
Interview Question
Why doesn’t standard CPU-based autoscaling work well for a self-hosted model on Kubernetes?
A model-serving pod’s real bottleneck is usually GPU utilization or request queue depth, not CPU — CPU can look idle while the GPU is saturated. Autoscaling needs a custom metric tied to actual inference load, and GPU-backed nodes are also a scarcer, more expensive resource than CPU nodes, so overprovisioning is costlier.
Explain It in 30 Seconds
Kubernetes runs and scales the containers behind an AI service — API, gateway, and any self-hosted model — but GPU-backed model-serving pods need GPU-aware scheduling and custom autoscaling metrics, since standard CPU-based autoscaling doesn’t reflect real inference load.
Real-World Stack
Technologies commonly used to implement this in production.