A request crosses three processes on its way to the GPU and back. This topic follows it across each boundary: what is sent, in what format, by which thread, and what can go wrong there. By the end you can say where any millisecond of a request’s latency was spent.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Processes and Wires | Which processes make up a server, how do they talk, and how many CPU cores do they need? |
| 02 | The Frontend: From JSON to Token IDs | What happens to a request before the engine sees it, and when is it refused? |
| 03 | AsyncLLM and the Way Back: Tokens to Text | How do token IDs become streamed text, and how are stop strings found? |
| 04 | The Engine Core Loop | What exactly is one engine step, and how does the GPU avoid waiting for Python? |
Read The Map first. The general theory of serving systems — queues, backpressure, timeouts — is in Serving Systems; this topic is how one specific engine implements it.