Pidoku
The Request Path

The Request Path

BasicTopic4 lessons~3h 15m

Lessons, in order

About this topic

A request crosses three processes on its way to the GPU and back. This topic follows it across each boundary: what is sent, in what format, by which thread, and what can go wrong there. By the end you can say where any millisecond of a request’s latency was spent.

#LessonThe question it answers
01Processes and WiresWhich processes make up a server, how do they talk, and how many CPU cores do they need?
02The Frontend: From JSON to Token IDsWhat happens to a request before the engine sees it, and when is it refused?
03AsyncLLM and the Way Back: Tokens to TextHow do token IDs become streamed text, and how are stop strings found?
04The Engine Core LoopWhat exactly is one engine step, and how does the GPU avoid waiting for Python?

Read The Map first. The general theory of serving systems — queues, backpressure, timeouts — is in Serving Systems; this topic is how one specific engine implements it.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom