Everything in the earlier topics is the same for every model and every request. This topic covers what changes when you turn something on: guessing tokens ahead, serving many fine-tunes from one copy of the weights, feeding images, extracting tool calls and reasoning, and running in fewer bits. Each lesson shows where the feature hooks into the scheduler, the cache and the model runner you already know, because that is what determines its cost.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Speculative Decoding | How can one forward pass yield several tokens, and when does that make things slower? |
| 02 | LoRA Adapters: Many Models in One | How do many fine-tunes share one base model, and what does max_loras really limit? |
| 03 | Multimodal Inputs | What does an image cost, and how does it pass through a text engine? |
| 04 | Tool Calling and Reasoning | How does one token stream become content, reasoning and tool_calls? |
| 05 | Quantization in vLLM | Which formats run fast on my GPU, and what do fewer bits buy a server? |