This topic builds an AI platform from the bottom up: what you need, in what order, starting from an application that calls someone else’s API and ending with your own GPU cluster serving many teams. Each lesson adds one layer and says when that layer is worth owning.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | The Layers and the Build-or-Buy Ladder | What does an AI platform consist of, and how much of it should I run? |
| 02 | Compute and Capacity | Which accelerators, how many, and bought how? |
| 03 | The Cluster | How do GPUs become a shared, schedulable resource? |
| 04 | The Serving Layer | How do weights become a fast, scalable endpoint? |
| 05 | The Data Layer | Where do weights, documents, vectors, state and events live? |
| 06 | Platform and Operations | How do many teams share it safely, and how is it run? |
The lessons stay at design level. For the mechanisms underneath — batching, KV caches, interconnects, GPU telemetry — each lesson links into Inference Engineering, GPU Engineering and Observability Engineering.