Pidoku

Compute and Capacity

Intermediate 50 min Difficulty 3/5 Lesson 02 of 06

Prerequisites The Layers and the Build-or-Buy Ladder, The Design Method

The idea in one minute#

Choosing compute for inference is three questions in order. Does the model fit? — memory decides which accelerators are even candidates. How many do I need? — measured throughput at your latency target, divided into peak demand, with headroom. How do I pay for them? — a base of committed capacity for the steady load and something elastic for the peaks. Get the order wrong and you buy the fastest chip that cannot hold your model, or forty GPUs for a load that needed six.

A picture#

flowchart LR
  M[":huggingface: <b>Model</b><br/><small>parameters, precision, context</small>"] --> FIT{":i-memory-stick: <b>Fits?</b><br/><small>weights + KV cache + overhead</small>"}
  FIT -->|"no"| OPT[":i-wrench: <b>Quantize, or split<br/>across GPUs</b>"]
  OPT --> FIT
  FIT -->|"yes"| THR[":i-gauge: <b>Measure throughput</b><br/><small>tokens/s at the TTFT target</small>"]
  D[":i-users: <b>Peak demand</b><br/><small>tokens/s from the ten numbers</small>"] --> N
  THR --> N[":i-cpu: <b>Replica count</b><br/><small>demand ÷ throughput ÷ occupancy<br/>+ spares</small>"]
  N --> BUY[":i-coins: <b>Purchase mix</b><br/><small>reserved base, on-demand peak,<br/>spot for batch</small>"]
  class M,D neutral
  class FIT,N queue
  class OPT warn
  class THR compute
  class BUY memory

How it really works#

Step 1 — does it fit?#

GPU memory holds three things, and people forget the second:

weights     parameters × bytes per parameter
            70 B parameters × 2 bytes (FP16)        = 140 GB
            70 B parameters × 1 byte  (FP8 / INT8)  =  70 GB
            70 B parameters × 0.5 byte (4-bit)      =  35 GB

KV cache    grows with concurrent sequences × context length
            often as large as the weights at production concurrency

overhead    activations, engine buffers, fragmentation: reserve ~10%

If the weights alone take 90% of the memory, the GPU can serve almost nobody at once. A model “fits” when there is room for the KV cache of your target concurrency. Memory planning for LLMs and KV cache math work through real cases.

When it does not fit: quantize (the cheapest fix, validated by your evaluation suite), choose an accelerator with more memory, or split the model across several GPUs with tensor parallelism — which works well inside one node and badly across slow networks.

Step 2 — what limits speed#

Generating tokens is limited by memory bandwidth, not arithmetic: each output token requires reading the model’s weights once. So when comparing accelerators for decoding, look at bandwidth before FLOPs.

What you see on a spec sheetWhat it tells an inference designer
Memory capacity (GB)Which models fit, and how much KV cache is left
Memory bandwidth (TB/s)The ceiling on single-stream token speed
Tensor arithmetic at low precisionPrefill speed and batch throughput
Interconnect (NVLink, InfiniBand, Ethernet)Whether multi-GPU and multi-node serving are practical
Power per GPU and cooling typeWhether your facility can host it at all

As of October 2026: Hopper (H100, H200) remains the largest installed base and the best-understood target; Blackwell is what most new capacity runs on; Rubin began shipping in August 2026, in liquid-cooled racks only. The current numbers are in GPU generations. Plan on what you can actually obtain, not on what was announced.

Step 3 — how many#

replicas = peak output tokens/s ÷ measured tokens/s per replica ÷ target occupancy
           + N+1 for failure  + capacity for rollouts (a new version runs beside the old)

“Measured” means a load test with your prompt and output lengths, at your TTFT and TPOT targets. A replica pushed to maximum throughput has terrible latency; the number that matters is throughput while still meeting the SLO — goodput. Expect that to be perhaps half of the benchmark headline.

Then size for the bad day: one replica lost, a deploy in progress and a traffic peak, at the same time. A design that is fine on the average day is not finished.

Step 4 — different workloads want different pools#

PoolWorkloadOptimise forTypical shape
InteractiveChat, copilots, agent steps a user is waiting onTTFT and TPOTHeadroom, fast scale-out, priority
BatchEvaluations, document processing, embeddings, overnight agentsTokens per dollarRun hot, queue, tolerate preemption
Small modelsRouters, classifiers, guardrails, embeddings, rerankersLatency at low costFractional or older GPUs, many replicas
ExperimentFine-tuning, evaluation of new modelsAvailability on demandQuotas, time limits, borrowed idle capacity

Sharing one pool across all four wastes money and hurts latency. Small models in particular should not occupy a whole flagship GPU each: sharing through MPS, MIG or time-slicing is covered in Sharing one GPU.

Step 5 — how to pay#

Purchase typePriceRiskUse for
Reserved / committedLowest per hourYou pay when idleThe load you have every hour of every day
On-demandHighestMay not be available when you need it mostPeaks, launches, failover
Spot / preemptibleLowCan vanish with short noticeBatch that checkpoints
Hosted API as overflowPer tokenDifferent model behaviourPeaks beyond your fleet, via the gateway

A dependable pattern: commit to the trough of your daily demand curve, cover the daily peak with autoscaled on-demand capacity, run batch on spot at night, and let the gateway overflow to an API when everything is full. GPU scarcity is real — an autoscaler cannot scale onto capacity the provider does not have — so test that the peak capacity can actually be obtained.

Capacity planning is a loop#

Estimates are wrong. What makes the plan work is the feedback: the gateway’s usage events give actual tokens per second by model and tenant; engine metrics give actual goodput per replica; the two together give occupancy, and occupancy over time is the input to the next purchase. Capacity planning covers the forecasting.

Code#

Sizing a fleet: does it fit, how many replicas, and what does the bad day need?

Go
// fleet.go — fit check and replica count for a self-hosted model.
package main

import (
	"fmt"
	"math"
)

type model struct {
	name          string
	paramsB       float64 // billions of parameters
	bytesPerParam float64
	kvGBPerSeq    float64 // KV cache per concurrent sequence at the working context length
}

type gpu struct {
	name     string
	memGB    float64
	hourCost float64
}

func main() {
	m := model{"70B, 8-bit", 70, 1.0, 1.6}
	g := gpu{"80 GB class", 80, 2.50}
	const (
		gpusPerReplica = 2      // tensor-parallel across two GPUs in one node
		peakOutTokSec  = 9000.0 // from the ten numbers
		goodputTokSec  = 1400.0 // measured per replica at the TTFT/TPOT target
		occupancy      = 0.6
	)

	mem := g.memGB * gpusPerReplica
	weights := m.paramsB * m.bytesPerParam
	free := mem*0.9 - weights // keep 10% for overhead
	fmt.Printf("%s on %d × %s: weights %.0f GB of %.0f GB\n", m.name, gpusPerReplica, g.name, weights, mem)
	if free <= 0 {
		fmt.Println("does not fit: quantize further or add GPUs per replica")
		return
	}
	fmt.Printf("room for KV cache: %.0f GB → about %.0f concurrent sequences per replica\n\n", free, math.Floor(free/m.kvGBPerSeq))

	base := math.Ceil(peakOutTokSec / goodputTokSec / occupancy)
	withSpare := base + 1                  // survive one replica failing
	rollout := math.Ceil(withSpare * 0.25) // a canary slice of the new version
	total := withSpare + rollout
	fmt.Printf("replicas for peak at %.0f%% occupancy: %.0f\n", occupancy*100, base)
	fmt.Printf("+1 for failure, +%.0f during a rollout:  %.0f replicas = %.0f GPUs\n", rollout, total, total*gpusPerReplica)
	fmt.Printf("cost at that size: $%.0f/day\n", total*gpusPerReplica*g.hourCost*24)
	fmt.Printf("\nthe same peak at 90%% occupancy needs %.0f replicas — and has no room for a burst.\n",
		math.Ceil(peakOutTokSec/goodputTokSec/0.9))
}

Remember this#

  • Fit first: weights plus KV cache plus overhead. A model that only just fits serves nobody.
  • Decode speed follows memory bandwidth.
  • Size from measured goodput at your SLO, at target occupancy, for the bad day.
  • Separate pools for interactive, batch, small models and experiments.
  • Commit to the trough, autoscale the peak, put batch on spot, overflow to an API.

Try it#

  1. Run fleet.go. Change the model to 16-bit (bytesPerParam 2.0). What breaks, and what are the two ways out?
  2. Halve goodputTokSec, as happens when prompts get longer. How does the bill change?
  3. Draw your service’s demand over 24 hours. Mark the level you would reserve.

Check yourself#

  1. Why is “the weights fit in memory” not enough?
  2. Why is benchmark throughput higher than the number you should size with?
  3. Which workloads should never share a pool with interactive traffic, and why?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom