Pidoku

The Cluster

Intermediate 50 min Difficulty 3/5 Lesson 03 of 06

Prerequisites Compute and Capacity

The idea in one minute#

A cluster turns a pile of GPU machines into one pool that many workloads share. Four jobs define it: expose each accelerator to the scheduler, place work on devices that suit it, queue and share when demand exceeds supply, and keep nodes healthy. In 2026 the default answer for inference and agents is Kubernetes, because the rest of the platform — gateways, sandboxes, data services — already lives there. Slurm remains common for large training jobs, and Ray for Python-native pipelines; the four jobs are the same whichever you choose.

A picture#

flowchart TB
  subgraph CP["Control plane"]
    direction LR
    API[":kubernetes: <b>API server</b>"] --> SCH[":kubernetes: <b>Scheduler</b><br/><small>DRA: match claims to devices</small>"]
    Q[":kueue: <b>Kueue</b><br/><small>quotas, queues, fair sharing</small>"] --> API
    AS[":keda: <b>Autoscalers</b><br/><small>replicas and nodes</small>"] --> API
  end
  subgraph NODE["GPU node"]
    direction LR
    KL[":kubernetes: <b>kubelet</b>"] --> RT[":containerd: <b>containerd</b><br/><small>+ NVIDIA container toolkit</small>"]
    DRV[":nvidia: <b>GPU Operator</b><br/><small>driver, DRA driver, DCGM</small>"] --> KL
    RT --> POD[":vllm: <b>Model server pod</b>"]
    POD --> GPU[(":nvidia: <b>GPUs</b>")]
    POD --> NV[(":i-hard-drive: <b>NVMe cache</b><br/><small>weights</small>")]
  end
  SCH --> KL
  DRV -.->|"publishes device attributes"| API
  REG[(":harbor: <b>Registry</b><br/><small>images and model artifacts</small>")] --> RT
  class API,SCH neutral
  class Q,AS queue
  class KL,RT neutral
  class DRV io
  class POD compute
  class GPU,NV,REG memory

How it really works#

Expose: from device plugin to DRA#

Kubernetes did not originally know what a GPU was. The device plugin taught it to count: a pod asks for nvidia.com/gpu: 1 and receives some whole GPU. That is still in wide use and is enough for simple cases.

Dynamic Resource Allocation (DRA) replaces counting with describing. Its core API has been stable since Kubernetes 1.34. A driver publishes each device with its attributes — model, memory, which NUMA node and PCIe root it hangs off — and a workload submits a claim stating what it needs:

a ResourceClaim, in words

  "one device from class gpu.nvidia.com
   where memory ≥ 80 GB and architecture is Blackwell,
   or failing that, two devices with ≥ 48 GB each"

What DRA makes possible that counting could not: choosing a GPU by attribute, sharing one device between containers, asking for alternatives in priority order, carving a GPU into partitions on demand, and tainting a single faulty device instead of a whole node. NVIDIA donated its DRA driver for GPUs to the Kubernetes community in 2026. New clusters should plan on DRA. The mechanics are in GPUs in Kubernetes and Kubernetes for GPUs.

Place: what a good placement respects#

  • Topology. GPUs that will run one model with tensor parallelism should be on the same node and the same fast interconnect.
  • Co-location of model and cache. A node that already has the weights on local NVMe starts a replica in seconds instead of minutes.
  • Groups. A multi-pod workload — a leader and workers serving one large model — must start together or not at all. Kubernetes added gang scheduling in 1.35 and split it into Workload and PodGroup APIs in 1.36, with preemption that evicts whole groups rather than stranding half of one.
  • Fragmentation. Eight free GPUs spread one per node cannot host an eight-GPU job. Bin-packing small work onto few nodes keeps large slots free.

Queue and share#

A GPU cluster is always oversubscribed in intent: every team would use more if it could. Kueue adds what the base scheduler lacks:

  • Quotas per team, in GPUs or in specific device classes.
  • Borrowing: a team may use another’s idle quota and gives it back when the owner needs it.
  • Queues with priority: a job that does not fit waits instead of failing.
  • Preemption rules: interactive serving outranks batch; batch outranks experiments.

Kueue is DRA-aware, so quotas can be expressed in the same device terms that claims use. The policy that works for most organisations: guaranteed quota for production serving, borrowable quota for batch, and a low-priority queue that soaks up whatever is idle.

Keep nodes healthy#

The NVIDIA GPU Operator installs and upgrades what a GPU node needs: driver, container toolkit, device plugin or DRA driver, and the DCGM exporter for telemetry. GPUs fail in ways CPUs rarely do — memory errors, a device falling off the bus, thermal throttling — so the cluster needs a loop that detects a bad device, taints it, drains work from it and opens a ticket. Without that loop, one faulty GPU serves errors for days. GPU telemetry covers what to watch.

The three things that surprise people#

Weights are large and loading is slow. A 100 GB model pulled over the network at start-up makes scale-out take ten minutes. Options: bake weights into node images, keep a node-local NVMe cache, or mount them as an OCI artifact — Kubernetes can now mount an image as a volume — so weights are pulled, cached and verified like any other image.

Node autoscaling is slow too. A new GPU node means provisioning, driver initialisation, image pull and model load. Keep warm spare capacity for interactive pools; do not expect an empty-to-serving time under several minutes.

The network matters once you leave one node. Serving one model across nodes, or moving KV cache between replicas, needs RDMA-class networking. Stay within a node where you can. See interconnects.

A reasonable cluster layout#

node pool            hardware                 runs                          scaling
─────────            ────────                 ────                          ───────
system               CPU                      control plane add-ons,        fixed
                                              gateways, operators
serving-interactive  flagship GPUs, NVMe      latency-critical models       warm spares + autoscale
serving-small        older or shared GPUs     embeddings, rerankers,        autoscale
                                              guardrail models
batch                any GPUs, spot allowed   evaluations, ingestion,       queue-driven, to zero
                                              offline agents
sandbox              CPU, isolated runtime    agent code execution          warm pool
data                 CPU, fast disks          vector index, databases       fixed + slow scale

Separate pools by trust as well as by hardware: sandboxes that run model-written code should not share nodes with the pods that hold credentials.

What fills the box in 2026#

JobOptions
OrchestratorKubernetes; Slurm for training-heavy sites; Ray on top of either
Device exposureNVIDIA GPU Operator with the DRA driver; device plugin on older clusters
Queueing and quotaKueue; Volcano; Run:ai-style commercial schedulers
AutoscalingKEDA or HPA on custom metrics for replicas; Karpenter or Cluster Autoscaler for nodes
ArtifactsAn OCI registry such as Harbor for images and model artifacts

Remember this#

  • Four jobs: expose, place, queue and share, keep healthy.
  • DRA describes devices by attribute; it is the direction for new clusters.
  • Quotas with borrowing and a batch queue are what make a shared cluster fair and full.
  • Model loading and node start-up are slow: cache weights locally and keep warm spares.
  • Pools are separated by hardware, by workload and by trust.

Try it#

  1. Write, in words, the claim for a workload needing four GPUs on one node with at least 80 GB each. What should happen if no node has four free?
  2. Three teams share 40 GPUs. Propose quotas, borrowing rules and priorities, then describe what happens when all three want 20 at once.
  3. Estimate cold-start time for a 60 GB model on a new node: provision, pull, load. Which term can you remove?

Check yourself#

  1. What can a DRA claim express that nvidia.com/gpu: 1 cannot?
  2. Why does fragmentation leave GPUs idle while jobs wait?
  3. Why should sandbox nodes be a separate pool?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom