Pidoku
Scaling and Operating

Scaling and Operating

ExpertTopic5 lessons~4h

Lessons, in order

About this topic

One engine on one GPU is the unit. This topic is about everything beyond it: spreading a model over several GPUs or running many copies, moving cached state between machines, reading the numbers a running server publishes, tuning by measurement instead of folklore, and running a fleet on Kubernetes. Each lesson leans on the internals from the earlier topics, because every operational question — why is it slow, how many do I need, what should I scale on — is answered by a mechanism you have already read.

#LessonThe question it answers
01Parallelism: TP, PP, DP, EP and CPHow should I lay a model out across my GPUs?
02Disaggregation and KV ConnectorsWhen is it worth moving the KV cache instead of recomputing it?
03Metrics: Reading a Running ServerWhat does each metric measure, and in what order do I read them?
04Tuning and BenchmarkingHow do I measure honestly and choose the dozen flags that matter?
05Deploying on KubernetesWhat must be configured differently from an ordinary web service?

Related material in the other courses: Distributed Inference, Production Inference, Observability for AI Infrastructure and AI Infrastructure From Scratch.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom