One engine on one GPU is the unit. This topic is about everything beyond it: spreading a model over several GPUs or running many copies, moving cached state between machines, reading the numbers a running server publishes, tuning by measurement instead of folklore, and running a fleet on Kubernetes. Each lesson leans on the internals from the earlier topics, because every operational question — why is it slow, how many do I need, what should I scale on — is answered by a mechanism you have already read.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Parallelism: TP, PP, DP, EP and CP | How should I lay a model out across my GPUs? |
| 02 | Disaggregation and KV Connectors | When is it worth moving the KV cache instead of recomputing it? |
| 03 | Metrics: Reading a Running Server | What does each metric measure, and in what order do I read them? |
| 04 | Tuning and Benchmarking | How do I measure honestly and choose the dozen flags that matter? |
| 05 | Deploying on Kubernetes | What must be configured differently from an ordinary web service? |
Related material in the other courses: Distributed Inference, Production Inference, Observability for AI Infrastructure and AI Infrastructure From Scratch.