Edition 01 AI infrastructure, from first principles
Learn modern AI infrastructure
Deep technical courses on how inference systems, GPUs, the telemetry that watches them and the language they are built in actually work — and how to design and secure the AI systems built on top — from first principles to production.
Courses
6 deep courses, organised by topic. Read one end to end, or jump straight to the lesson you need.
Inference Engineering
How modern LLM inference works — from a single matrix multiply to a multi-tenant serving platform.
02GPU Engineering
The GPU itself — what is inside the chip, how it is programmed, why it is fast, and how thousands are run together.
03Observability Engineering
Seeing inside running systems — metrics, logs, traces and profiles, then the GPU, token and cost telemetry that AI infrastructure adds.
04Golang Engineering
Go from the first program to the runtime — how values sit in memory, how the allocator, garbage collector and scheduler work, and how to build AI systems in it.
05AI System Design
Designing AI systems in 2026 — the building blocks, AI infrastructure from an empty rack to a platform, agentic systems, and full worked designs with the real tools named.
06AI Security Engineering
Securing AI applications at every stage — data, models and supply chain, build, deployment, runtime and agents — with the attacks explained first and the defences that actually hold.
New courses are added as plain Markdown folders. Distributed systems for AI and networking for GPU clusters are on the list.
Five levels, one order
Every course runs through the same five levels. Each level assumes the ones before it and nothing else.
- 01FoundationsBuild the mental model.
- 02BasicUnderstand the core mechanisms.
- 03IntermediateLearn the optimization techniques.
- 04AdvancedStudy systems at production scale.
- 05ExpertDesign platforms and read the frontier.
- Start with fundamentals. No step assumes knowledge the course has not taught.
- Understand the system. Mechanisms are built in small Go programs you can run.
- Study the bottlenecks. Memory, bandwidth, queues — find what actually limits speed.
- Learn the optimization techniques. Each one is derived from the bottleneck it removes.
- Build real systems. Every course ends in projects, not quizzes.