The idea in one minute#
Choosing compute for inference is three questions in order. Does the model fit? — memory decides which accelerators are even candidates. How many do I need? — measured throughput at your latency target, divided into peak demand, with headroom. How do I pay for them? — a base of committed capacity for the steady load and something elastic for the peaks. Get the order wrong and you buy the fastest chip that cannot hold your model, or forty GPUs for a load that needed six.
A picture#
flowchart LR
M[":huggingface: <b>Model</b><br/><small>parameters, precision, context</small>"] --> FIT{":i-memory-stick: <b>Fits?</b><br/><small>weights + KV cache + overhead</small>"}
FIT -->|"no"| OPT[":i-wrench: <b>Quantize, or split<br/>across GPUs</b>"]
OPT --> FIT
FIT -->|"yes"| THR[":i-gauge: <b>Measure throughput</b><br/><small>tokens/s at the TTFT target</small>"]
D[":i-users: <b>Peak demand</b><br/><small>tokens/s from the ten numbers</small>"] --> N
THR --> N[":i-cpu: <b>Replica count</b><br/><small>demand ÷ throughput ÷ occupancy<br/>+ spares</small>"]
N --> BUY[":i-coins: <b>Purchase mix</b><br/><small>reserved base, on-demand peak,<br/>spot for batch</small>"]
class M,D neutral
class FIT,N queue
class OPT warn
class THR compute
class BUY memoryHow it really works#
Step 1 — does it fit?#
GPU memory holds three things, and people forget the second:
weights parameters × bytes per parameter
70 B parameters × 2 bytes (FP16) = 140 GB
70 B parameters × 1 byte (FP8 / INT8) = 70 GB
70 B parameters × 0.5 byte (4-bit) = 35 GB
KV cache grows with concurrent sequences × context length
often as large as the weights at production concurrency
overhead activations, engine buffers, fragmentation: reserve ~10%If the weights alone take 90% of the memory, the GPU can serve almost nobody at once. A model “fits” when there is room for the KV cache of your target concurrency. Memory planning for LLMs and KV cache math work through real cases.
When it does not fit: quantize (the cheapest fix, validated by your evaluation suite), choose an accelerator with more memory, or split the model across several GPUs with tensor parallelism — which works well inside one node and badly across slow networks.
Step 2 — what limits speed#
Generating tokens is limited by memory bandwidth, not arithmetic: each output token requires reading the model’s weights once. So when comparing accelerators for decoding, look at bandwidth before FLOPs.
| What you see on a spec sheet | What it tells an inference designer |
|---|---|
| Memory capacity (GB) | Which models fit, and how much KV cache is left |
| Memory bandwidth (TB/s) | The ceiling on single-stream token speed |
| Tensor arithmetic at low precision | Prefill speed and batch throughput |
| Interconnect (NVLink, InfiniBand, Ethernet) | Whether multi-GPU and multi-node serving are practical |
| Power per GPU and cooling type | Whether your facility can host it at all |
As of October 2026: Hopper (H100, H200) remains the largest installed base and the best-understood target; Blackwell is what most new capacity runs on; Rubin began shipping in August 2026, in liquid-cooled racks only. The current numbers are in GPU generations. Plan on what you can actually obtain, not on what was announced.
Step 3 — how many#
replicas = peak output tokens/s ÷ measured tokens/s per replica ÷ target occupancy
+ N+1 for failure + capacity for rollouts (a new version runs beside the old)“Measured” means a load test with your prompt and output lengths, at your TTFT and TPOT targets. A replica pushed to maximum throughput has terrible latency; the number that matters is throughput while still meeting the SLO — goodput. Expect that to be perhaps half of the benchmark headline.
Then size for the bad day: one replica lost, a deploy in progress and a traffic peak, at the same time. A design that is fine on the average day is not finished.
Step 4 — different workloads want different pools#
| Pool | Workload | Optimise for | Typical shape |
|---|---|---|---|
| Interactive | Chat, copilots, agent steps a user is waiting on | TTFT and TPOT | Headroom, fast scale-out, priority |
| Batch | Evaluations, document processing, embeddings, overnight agents | Tokens per dollar | Run hot, queue, tolerate preemption |
| Small models | Routers, classifiers, guardrails, embeddings, rerankers | Latency at low cost | Fractional or older GPUs, many replicas |
| Experiment | Fine-tuning, evaluation of new models | Availability on demand | Quotas, time limits, borrowed idle capacity |
Sharing one pool across all four wastes money and hurts latency. Small models in particular should not occupy a whole flagship GPU each: sharing through MPS, MIG or time-slicing is covered in Sharing one GPU.
Step 5 — how to pay#
| Purchase type | Price | Risk | Use for |
|---|---|---|---|
| Reserved / committed | Lowest per hour | You pay when idle | The load you have every hour of every day |
| On-demand | Highest | May not be available when you need it most | Peaks, launches, failover |
| Spot / preemptible | Low | Can vanish with short notice | Batch that checkpoints |
| Hosted API as overflow | Per token | Different model behaviour | Peaks beyond your fleet, via the gateway |
A dependable pattern: commit to the trough of your daily demand curve, cover the daily peak with autoscaled on-demand capacity, run batch on spot at night, and let the gateway overflow to an API when everything is full. GPU scarcity is real — an autoscaler cannot scale onto capacity the provider does not have — so test that the peak capacity can actually be obtained.
Capacity planning is a loop#
Estimates are wrong. What makes the plan work is the feedback: the gateway’s usage events give actual tokens per second by model and tenant; engine metrics give actual goodput per replica; the two together give occupancy, and occupancy over time is the input to the next purchase. Capacity planning covers the forecasting.
Code#
Sizing a fleet: does it fit, how many replicas, and what does the bad day need?
// fleet.go — fit check and replica count for a self-hosted model.
package main
import (
"fmt"
"math"
)
type model struct {
name string
paramsB float64 // billions of parameters
bytesPerParam float64
kvGBPerSeq float64 // KV cache per concurrent sequence at the working context length
}
type gpu struct {
name string
memGB float64
hourCost float64
}
func main() {
m := model{"70B, 8-bit", 70, 1.0, 1.6}
g := gpu{"80 GB class", 80, 2.50}
const (
gpusPerReplica = 2 // tensor-parallel across two GPUs in one node
peakOutTokSec = 9000.0 // from the ten numbers
goodputTokSec = 1400.0 // measured per replica at the TTFT/TPOT target
occupancy = 0.6
)
mem := g.memGB * gpusPerReplica
weights := m.paramsB * m.bytesPerParam
free := mem*0.9 - weights // keep 10% for overhead
fmt.Printf("%s on %d × %s: weights %.0f GB of %.0f GB\n", m.name, gpusPerReplica, g.name, weights, mem)
if free <= 0 {
fmt.Println("does not fit: quantize further or add GPUs per replica")
return
}
fmt.Printf("room for KV cache: %.0f GB → about %.0f concurrent sequences per replica\n\n", free, math.Floor(free/m.kvGBPerSeq))
base := math.Ceil(peakOutTokSec / goodputTokSec / occupancy)
withSpare := base + 1 // survive one replica failing
rollout := math.Ceil(withSpare * 0.25) // a canary slice of the new version
total := withSpare + rollout
fmt.Printf("replicas for peak at %.0f%% occupancy: %.0f\n", occupancy*100, base)
fmt.Printf("+1 for failure, +%.0f during a rollout: %.0f replicas = %.0f GPUs\n", rollout, total, total*gpusPerReplica)
fmt.Printf("cost at that size: $%.0f/day\n", total*gpusPerReplica*g.hourCost*24)
fmt.Printf("\nthe same peak at 90%% occupancy needs %.0f replicas — and has no room for a burst.\n",
math.Ceil(peakOutTokSec/goodputTokSec/0.9))
}Remember this#
- Fit first: weights plus KV cache plus overhead. A model that only just fits serves nobody.
- Decode speed follows memory bandwidth.
- Size from measured goodput at your SLO, at target occupancy, for the bad day.
- Separate pools for interactive, batch, small models and experiments.
- Commit to the trough, autoscale the peak, put batch on spot, overflow to an API.
Try it#
- Run
fleet.go. Change the model to 16-bit (bytesPerParam2.0). What breaks, and what are the two ways out? - Halve
goodputTokSec, as happens when prompts get longer. How does the bill change? - Draw your service’s demand over 24 hours. Mark the level you would reserve.
Check yourself#
- Why is “the weights fit in memory” not enough?
- Why is benchmark throughput higher than the number you should size with?
- Which workloads should never share a pool with interactive traffic, and why?