The idea in one minute#
A model’s weights are normally stored as 16-bit numbers. Quantization stores them in 8 or 4 bits. In a serving engine that buys two different things. Decode speed is limited by how fast the GPU can read the weights, so half the bytes is close to twice the speed. And the memory freed from weights becomes KV cache blocks, so the same GPU holds more concurrent requests. vLLM does not quantize models itself in the usual case; it loads models quantized by other tools and runs them with kernels matched to each format. It can also quantize the KV cache independently of the weights. Which formats work on which hardware, and which actually run fast, is a table worth knowing before you choose.
A picture#
flowchart LR FP[":i-hard-drive: <b>16-bit model</b><br/><small>15 GiB</small>"] --> T[":i-wrench: <b>Quantization tool</b><br/><small>LLM Compressor, AutoAWQ,<br/>GPTQModel, ModelOpt…</small>"] T --> Q[":i-archive: <b>Quantized checkpoint</b><br/><small>+ a config naming the scheme</small>"] Q --> LD[":vllm: <b>vLLM loader</b><br/><small>reads the config,<br/>builds quantized layers</small>"] LD --> K[":nvidia: <b>Quantized linear kernel</b><br/><small>Marlin, CUTLASS, …</small>"] LD --> MEM["weights: 4–8 GiB<br/><small>the rest becomes KV blocks</small>"] ONL[":i-zap: <b>Online quantization</b><br/><small>--quantization fp8_per_tensor</small>"] -.->|"no pre-quantized checkpoint"| LD KV[":i-layers: <b>--kv-cache-dtype fp8</b><br/><small>independent of the weights</small>"] -.-> MEM class FP,Q memory class T,LD compute class K compute class MEM neutral class ONL,KV queue
How it really works#
Two benefits, one of them specific to serving#
For the theory of the methods, see Quantization Overview. What matters inside vLLM is the arithmetic of Sizing the Cache:
KV cache memory = total × 0.92 − weights − working memoryWeights are the largest term being subtracted. On a 24 GiB GPU with an 8B model:
16-bit weights 15 GiB → about 6 GiB of KV cache
8-bit weights 7.5 GiB → about 14 GiB of KV cache 2.2× the tokens in flight
4-bit weights 4 GiB → about 17 GiB of KV cache 2.8×So quantization raises a server’s concurrency more than proportionally, in addition to raising per-request speed. That second effect is invisible in a single-request benchmark and is often the larger one in production.
How vLLM knows a model is quantized#
You normally pass nothing. A quantized checkpoint carries a configuration describing its
scheme, and each supported format is a QuantizationConfig class that declares which config
files it recognises, the minimum GPU capability it needs, and — per layer — which quantized
layer implementation to build:
| Method | Role |
|---|---|
get_config_filenames() | Files to look for in the model directory |
from_config(config) | Build the scheme from that file |
get_min_capability() | Minimum GPU compute capability, e.g. 80 for Ampere |
get_quant_method(layer, prefix) | The implementation for a given layer, or None to leave it unquantized |
That last method is why mixed-precision models work: different layers can return different
methods. It relies on the prefix argument every vLLM model constructor takes
(Executor, Worker, Model Runner), and on
the rule that layers are built quantized from the start rather than converted after loading.
--quantization <scheme> forces a scheme when detection is ambiguous, and selects the
online schemes described below.
Formats#
| Format | Bits (weights / activations) | Produced by | Notes |
|---|---|---|---|
| FP8 W8A8 | 8 / 8 | LLM Compressor | Small quality loss; needs a GPU with native 8-bit float support |
| INT8 W8A8 | 8 / 8 | LLM Compressor | Runs on older GPUs and on CPUs |
| INT4 W4A16 | 4 / 16 | LLM Compressor, AutoAWQ, GPTQModel | The common “4-bit” format; weights only |
| AWQ, GPTQ | 4 (typically) / 16 | AutoAWQ, GPTQModel | Calibrated weight-only methods |
| NVFP4, MXFP4, MXFP8 | 4 or 8 / varies | NVIDIA Model Optimizer, LLM Compressor | Block-scaled formats for the newest GPUs |
| bitsandbytes | 4 or 8 | Loaded on the fly | Convenient; not the fastest path |
| GGUF | various | llama.cpp ecosystem | Supported for compatibility; limited and slower than native formats |
LLM Compressor is the vLLM project’s own quantization library and the recommended starting point: its outputs are the formats vLLM’s fast paths are built for.
Hardware support is uneven#
From the documentation’s compatibility table:
| Implementation | Turing | Ampere | Ada | Hopper | AMD GPU | x86 CPU |
|---|---|---|---|---|---|---|
| AWQ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ |
| GPTQ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ |
| Marlin kernels (GPTQ/AWQ/FP8/FP4) | ✅* | ✅ | ✅ | ✅ | ❌ | ❌ |
| LLM Compressor INT8 (W8A8) | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ |
| LLM Compressor FP8 (W8A8) | ❌ | ❌ | ✅ | ✅ | ✅ | ❌ |
| bitsandbytes | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ |
| GGUF | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ |
Two lines to read carefully. FP8 W8A8 needs Ada or Hopper generation and newer: an A100 (Ampere) cannot run it natively. And Marlin is not a format but a family of fast kernels for 4-bit and 8-bit weights; a format that runs through Marlin is fast, and one that does not can be slower than the 16-bit model despite being smaller.
That is the general warning about quantization in any engine: a smaller file is not automatically faster. Speed comes from a kernel that computes directly on the packed weights. Without one, weights are unpacked to 16 bits on every use, and you have saved memory but not time.
Online quantization#
For a model with no quantized checkpoint, vLLM can quantize while loading:
vllm serve meta-llama/Llama-3.1-8B --quantization fp8_per_tensorThe documentation describes it: “take a BF16/FP16 model and quantize its Linear and MoE weights
to lower precision (such as FP8) at load time, without needing a pre-quantized checkpoint or
calibration data. Weights are converted during model loading and activations are dynamically
scaled during each forward pass.” Schemes include fp8_per_tensor, fp8_per_block, mxfp8
and an MXFP4 variant.
It is the quickest way to try a quantized deployment. A calibrated, pre-quantized checkpoint is usually more accurate, and loads faster because nothing is converted at startup.
The KV cache can be quantized separately#
--kv-cache-dtype fp8 stores cached keys and values in 8 bits whatever the weights are. It
halves bytes per token and therefore doubles the pool, as the calculator in
Sizing the Cache showed.
How the 16-bit values are scaled into 8 bits matters for quality:
| Approach | Scales | Accuracy |
|---|---|---|
| No calibration | All 1.0 | Works; least accurate |
| Per-tensor calibration | One scale each for Q, K and V | Better |
| Per-attention-head calibration | One scale per head | Best; FlashAttention backend only, calibrated with LLM Compressor |
Details that interact with earlier lessons:
- The attention backend must support it. An FP8 cache changes which backends validate and their priority order (Attention Backends). With some backends attention itself is then computed in 8 bits.
- Some layers are more sensitive.
--kv-cache-dtype-skip-layersleaves chosen layers (by index or by type, such as sliding-window layers) at full precision. - Some models default to it. For certain latent-attention models the default cache type is
already FP8, and you opt out with
--kv-cache-dtype bfloat16. - Cached blocks are not portable across the setting. A KV cache moved between instances (Disaggregation and KV Connectors) must be produced and consumed with the same cache type.
What does not change#
Quantization does not alter the scheduler, the block pool’s logic, or the prefix cache’s behaviour. Blocks are still 16 tokens; hashes are still over token IDs. Only the number of blocks and the time per step change. This is why it composes freely with every scheduling feature in this course.
Choosing#
| Situation | Reasonable choice |
|---|---|
| Hopper/Ada or newer, quality-sensitive | FP8 W8A8 from LLM Compressor |
| Ampere (A100, A10G) | INT4 W4A16 or AWQ through Marlin; or INT8 W8A8 |
| Model does not fit at all | 4-bit weights, then measure quality on your own task |
| Model fits, cache is the constraint | Keep weights as they are; --kv-cache-dtype fp8 |
| No quantized checkpoint, want to try | Online fp8_per_tensor |
| CPU serving | INT8 W8A8 |
Always evaluate on your own workload. Aggregate benchmark scores hide task-specific losses, and 4-bit models in particular degrade unevenly: fluent prose survives, exact arithmetic and long structured outputs suffer first.
Code#
The serving effect in numbers: for one GPU and one model, how weight precision and cache precision change capacity and decode speed.
package main
import "fmt"
const GiB = 1 << 30
func main() {
const (
gpuGiB = 24.0
util = 0.92
otherGiB = 0.8 // activations, CUDA graphs, non-torch
params = 8.03 // billions of parameters
kvPerTok16 = 131072.0
avgReqTokens = 3000.0 // prompt + output of a typical request
bandwidthGBs = 700.0 // GPU memory bandwidth, GB/s (order of magnitude)
)
type cfg struct {
name string
weightBits float64
kvBits float64
fastKernel bool // false: weights are unpacked to 16 bits on every use
}
cfgs := []cfg{
{"16-bit weights, 16-bit KV", 16, 16, true},
{"16-bit weights, 8-bit KV", 16, 8, true},
{" 8-bit weights, 16-bit KV", 8, 16, true},
{" 8-bit weights, 8-bit KV", 8, 8, true},
{" 4-bit weights, 16-bit KV", 4, 16, true},
{" 4-bit weights, 8-bit KV", 4, 8, true},
{" 4-bit, no native kernel ", 4, 16, false},
}
fmt.Printf("%-28s %9s %9s %12s %12s %14s\n",
"configuration", "weights", "KV pool", "tokens", "requests", "decode steps/s")
for _, c := range cfgs {
weights := params * 1e9 * c.weightBits / 8 / GiB
kv := gpuGiB*util - weights - otherGiB
tokens := kv * GiB / (kvPerTok16 * c.kvBits / 16)
// A decode step must read every weight once: steps/s is bounded by bandwidth.
readGB := params * c.weightBits / 8
if !c.fastKernel {
readGB = params*c.weightBits/8 + params*2 // read packed, then touch the 16-bit copy
}
fmt.Printf("%-28s %6.1f GiB %6.1f GiB %12.0f %12.0f %14.0f\n",
c.name, weights, kv, tokens, tokens/avgReqTokens, bandwidthGBs/readGB)
}
}Three things to take from the table. Moving from 16-bit to 4-bit weights with a native kernel roughly quadruples the decode rate bound and nearly triples concurrent requests. Quantizing only the cache doubles concurrency without touching the weights. And the last row — 4-bit weights without a kernel that computes on them — has all the memory benefit and is slower than the 16-bit model.
Remember this#
- Quantization helps a server twice: faster decode, and more memory left for KV blocks.
- vLLM loads models quantized by other tools; the checkpoint’s config selects the scheme.
- LLM Compressor is the project’s own tool and targets vLLM’s fast paths.
- FP8 weights need Ada/Hopper-class GPUs or newer; 4-bit through Marlin runs on Ampere too.
- A format without a native kernel saves memory but can be slower.
- Online quantization converts a 16-bit model at load time with no calibration.
--kv-cache-dtype fp8doubles cache capacity independently of the weights; calibrated scales improve its accuracy.- The scheduler and cache logic are unchanged; only block count and step time move.
Try it#
- Change the GPU to 80 GiB and the model to 70 billion parameters. Which configurations fit on one GPU at all?
- With
avgReqTokensat 30,000 (long documents), which matters more for concurrency: weight precision or cache precision? - On a real server, compare
Available KV cache memoryand tokens per second for the same model at 16 bits and with--quantization fp8_per_tensor.
Check yourself#
- Why does quantizing weights increase the number of requests a server can hold?
- An INT4 model is smaller but slower than the 16-bit original on your GPU. What is the likely reason?
- What does
--kv-cache-dtype fp8change, and what does it leave unchanged?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
- Quantization — formats and the hardware compatibility table
- Quantized KV cache
- Online quantization
- LLM Compressor
vllm/config/cache.py—cache_dtype