INFRA Signal 93
Kubernetes AI factory pattern shares GPU pools across teams with DRA scheduling and tenant isolation
A CNCF-authored post lays out a layered Kubernetes stack for running multi-tenant AI workloads on shared GPU fleets, combining Dynamic Resource Allocation, virtual clusters, and a broad set of CNCF and NVIDIA tools to keep utilization high while isolating teams.
GPU capital expense dominates AI infrastructure, and the traditional Kubernetes device-plugin model pins a whole accelerator per pod even at low utilization, stranding capacity. The post argues the real engineering problem is safe multi-tenant access to the same expensive hardware, not model serving throughput, and maps a concrete stack to solve it.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Dynamic Resource Allocation, GA in Kubernetes 1.34, lets the scheduler treat accelerators as rich devices with attributes, memory, and topology, though density still depends on the device layer underneath.
SemiAnalysis's ClusterMAX scoring rewards hard per-tenant isolation down to per-tenant clusters and DPU-based boundaries, while flagging weak boundaries like putting many tenants on one cluster.
The proposed stack spans bare-metal provisioning through billing, combining Metal3, Cluster API, vCluster, DRA, MIG, KAI Scheduler, vLLM, KServe, Cilium, Prometheus, Kyverno, and OpenCost among other tools.
THE READ
What the cluster adds up to.
The central claim is that an AI factory is a shared GPU pool serving multiple teams simultaneously, one fine-tuning, another serving inference, a third running evaluations, rather than a single model or cluster. The bottleneck is utilization of expensive accelerators, not peak tokens-per-second from a single run. SemiAnalysis's ClusterMAX scoring reflects this shift by grading GPU cloud providers on security, networking, storage, reliability, and support rather than raw throughput, and its security criteria reward hard per-tenant isolation down to per-tenant Kubernetes clusters and DPU-based isolation. The wrapper around the GPUs, not the GPUs themselves, is what gets judged.
Two structural problems keep utilization low. The resource model is the first: in the traditional device-plugin model a pod requests nvidia.com/gpu: 1 and pins a whole accelerator even at ten percent use. Dynamic Resource Allocation, GA in Kubernetes 1.34, lets the scheduler treat accelerators as rich devices with attributes, memory, and topology, but it does not by itself carve a GPU into fractions, density comes from the device layer underneath, such as MIG or time-slicing. The isolation model is the second: platforms default to a dedicated cluster or dedicated GPU block per team, which is safe when trust is strict but wastes most of the hardware. Operators in the field resort to bare-metal provisioners and manual workarounds, or hand each customer a dedicated block of GPUs and turn away demand they cannot isolate cleanly.
The proposed solution is a layered stack, most of it Kubernetes-native or CNCF projects, with a few OSS tools from NVIDIA. At the bottom, Metal3, Ironic, Tinkerbell, Redfish, and NetBox handle bare-metal provisioning and validation. Cluster API with Argo CD or Flux manages cluster lifecycle via GitOps, while Node Feature Discovery and the GPU and Network Operators from NVIDIA and AMD label nodes with GPU, NIC, MIG, and topology information. Tenant isolation uses vCluster for tenant clusters and sandboxed runtimes, keeping teams apart on the same hardware rather than dedicating clusters per team.
Above the isolation layer, GPU allocation combines DRA, MIG, HAMi, and time-slicing for device sharing, with KAI Scheduler providing Topology Aware Scheduling alongside Volcano and Kueue. Inference and serving run through vLLM, KServe, and llm-d, while Slinky brings SLURM workloads onto Kubernetes and KubeVirt handles VMs for tenants. Networking relies on Cilium, Multus, SR-IOV, and RDMA-based solutions to move data between GPUs and isolate tenants. The remaining layers, storage via CSI and Rook/Ceph, observability via Prometheus and DCGM exporter, identity and policy via Keycloak and Kyverno or OPA, secrets via OpenBao and External Secrets, reliability via DCGM health checks and Node Problem Detector, and billing via OpenCost and DCGM GPU-seconds, round out the stack into something that can provision, schedule, isolate, observe, and charge for GPU capacity.
The post does not claim any single layer solves the problem. The fix is not a new model server but a stack that allocates accelerators so capacity is neither stranded nor unsafe, and isolates tenants so packing them together holds up. The article is a single-source CNCF post by a platform advocate, not an independent benchmark or case study, so the stack should be read as a reference architecture assembled from available tools rather than a proven production pattern. What it does concretely establish is that the cloud-native ecosystem now supplies most of the parts needed to close the gap between Kubernetes' mature container primitives and the accelerator-sharing and tenant-isolation requirements of multi-team AI workloads.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗