Pidoku
Features

Features

AdvancedTopic5 lessons~3h 35m

Lessons, in order

About this topic

Everything in the earlier topics is the same for every model and every request. This topic covers what changes when you turn something on: guessing tokens ahead, serving many fine-tunes from one copy of the weights, feeding images, extracting tool calls and reasoning, and running in fewer bits. Each lesson shows where the feature hooks into the scheduler, the cache and the model runner you already know, because that is what determines its cost.

#LessonThe question it answers
01Speculative DecodingHow can one forward pass yield several tokens, and when does that make things slower?
02LoRA Adapters: Many Models in OneHow do many fine-tunes share one base model, and what does max_loras really limit?
03Multimodal InputsWhat does an image cost, and how does it pass through a text engine?
04Tool Calling and ReasoningHow does one token stream become content, reasoning and tool_calls?
05Quantization in vLLMWhich formats run fast on my GPU, and what do fewer bits buy a server?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom