Pidoku

Quantization in vLLM

Advanced 40 min Difficulty 3/5 Lesson 05 of 05

Prerequisites Sizing the Cache, Executor, Worker, Model Runner

The idea in one minute#

A model’s weights are normally stored as 16-bit numbers. Quantization stores them in 8 or 4 bits. In a serving engine that buys two different things. Decode speed is limited by how fast the GPU can read the weights, so half the bytes is close to twice the speed. And the memory freed from weights becomes KV cache blocks, so the same GPU holds more concurrent requests. vLLM does not quantize models itself in the usual case; it loads models quantized by other tools and runs them with kernels matched to each format. It can also quantize the KV cache independently of the weights. Which formats work on which hardware, and which actually run fast, is a table worth knowing before you choose.

A picture#

flowchart LR
  FP[":i-hard-drive: <b>16-bit model</b><br/><small>15 GiB</small>"] --> T[":i-wrench: <b>Quantization tool</b><br/><small>LLM Compressor, AutoAWQ,<br/>GPTQModel, ModelOpt…</small>"]
  T --> Q[":i-archive: <b>Quantized checkpoint</b><br/><small>+ a config naming the scheme</small>"]
  Q --> LD[":vllm: <b>vLLM loader</b><br/><small>reads the config,<br/>builds quantized layers</small>"]
  LD --> K[":nvidia: <b>Quantized linear kernel</b><br/><small>Marlin, CUTLASS, …</small>"]
  LD --> MEM["weights: 4–8 GiB<br/><small>the rest becomes KV blocks</small>"]
  ONL[":i-zap: <b>Online quantization</b><br/><small>--quantization fp8_per_tensor</small>"] -.->|"no pre-quantized checkpoint"| LD
  KV[":i-layers: <b>--kv-cache-dtype fp8</b><br/><small>independent of the weights</small>"] -.-> MEM
  class FP,Q memory
  class T,LD compute
  class K compute
  class MEM neutral
  class ONL,KV queue

How it really works#

Two benefits, one of them specific to serving#

For the theory of the methods, see Quantization Overview. What matters inside vLLM is the arithmetic of Sizing the Cache:

KV cache memory = total × 0.92 − weights − working memory

Weights are the largest term being subtracted. On a 24 GiB GPU with an 8B model:

16-bit weights   15 GiB  →  about  6 GiB of KV cache
 8-bit weights  7.5 GiB  →  about 14 GiB of KV cache      2.2× the tokens in flight
 4-bit weights    4 GiB  →  about 17 GiB of KV cache      2.8×

So quantization raises a server’s concurrency more than proportionally, in addition to raising per-request speed. That second effect is invisible in a single-request benchmark and is often the larger one in production.

How vLLM knows a model is quantized#

You normally pass nothing. A quantized checkpoint carries a configuration describing its scheme, and each supported format is a QuantizationConfig class that declares which config files it recognises, the minimum GPU capability it needs, and — per layer — which quantized layer implementation to build:

MethodRole
get_config_filenames()Files to look for in the model directory
from_config(config)Build the scheme from that file
get_min_capability()Minimum GPU compute capability, e.g. 80 for Ampere
get_quant_method(layer, prefix)The implementation for a given layer, or None to leave it unquantized

That last method is why mixed-precision models work: different layers can return different methods. It relies on the prefix argument every vLLM model constructor takes (Executor, Worker, Model Runner), and on the rule that layers are built quantized from the start rather than converted after loading.

--quantization <scheme> forces a scheme when detection is ambiguous, and selects the online schemes described below.

Formats#

FormatBits (weights / activations)Produced byNotes
FP8 W8A88 / 8LLM CompressorSmall quality loss; needs a GPU with native 8-bit float support
INT8 W8A88 / 8LLM CompressorRuns on older GPUs and on CPUs
INT4 W4A164 / 16LLM Compressor, AutoAWQ, GPTQModelThe common “4-bit” format; weights only
AWQ, GPTQ4 (typically) / 16AutoAWQ, GPTQModelCalibrated weight-only methods
NVFP4, MXFP4, MXFP84 or 8 / variesNVIDIA Model Optimizer, LLM CompressorBlock-scaled formats for the newest GPUs
bitsandbytes4 or 8Loaded on the flyConvenient; not the fastest path
GGUFvariousllama.cpp ecosystemSupported for compatibility; limited and slower than native formats

LLM Compressor is the vLLM project’s own quantization library and the recommended starting point: its outputs are the formats vLLM’s fast paths are built for.

Hardware support is uneven#

From the documentation’s compatibility table:

ImplementationTuringAmpereAdaHopperAMD GPUx86 CPU
AWQ✅✅✅✅❌✅
GPTQ✅✅✅✅❌✅
Marlin kernels (GPTQ/AWQ/FP8/FP4)✅*✅✅✅❌❌
LLM Compressor INT8 (W8A8)✅✅✅✅❌✅
LLM Compressor FP8 (W8A8)❌❌✅✅✅❌
bitsandbytes✅✅✅✅❌❌
GGUF✅✅✅✅✅❌

Two lines to read carefully. FP8 W8A8 needs Ada or Hopper generation and newer: an A100 (Ampere) cannot run it natively. And Marlin is not a format but a family of fast kernels for 4-bit and 8-bit weights; a format that runs through Marlin is fast, and one that does not can be slower than the 16-bit model despite being smaller.

That is the general warning about quantization in any engine: a smaller file is not automatically faster. Speed comes from a kernel that computes directly on the packed weights. Without one, weights are unpacked to 16 bits on every use, and you have saved memory but not time.

Online quantization#

For a model with no quantized checkpoint, vLLM can quantize while loading:

Shell
vllm serve meta-llama/Llama-3.1-8B --quantization fp8_per_tensor

The documentation describes it: “take a BF16/FP16 model and quantize its Linear and MoE weights to lower precision (such as FP8) at load time, without needing a pre-quantized checkpoint or calibration data. Weights are converted during model loading and activations are dynamically scaled during each forward pass.” Schemes include fp8_per_tensor, fp8_per_block, mxfp8 and an MXFP4 variant.

It is the quickest way to try a quantized deployment. A calibrated, pre-quantized checkpoint is usually more accurate, and loads faster because nothing is converted at startup.

The KV cache can be quantized separately#

--kv-cache-dtype fp8 stores cached keys and values in 8 bits whatever the weights are. It halves bytes per token and therefore doubles the pool, as the calculator in Sizing the Cache showed.

How the 16-bit values are scaled into 8 bits matters for quality:

ApproachScalesAccuracy
No calibrationAll 1.0Works; least accurate
Per-tensor calibrationOne scale each for Q, K and VBetter
Per-attention-head calibrationOne scale per headBest; FlashAttention backend only, calibrated with LLM Compressor

Details that interact with earlier lessons:

  • The attention backend must support it. An FP8 cache changes which backends validate and their priority order (Attention Backends). With some backends attention itself is then computed in 8 bits.
  • Some layers are more sensitive. --kv-cache-dtype-skip-layers leaves chosen layers (by index or by type, such as sliding-window layers) at full precision.
  • Some models default to it. For certain latent-attention models the default cache type is already FP8, and you opt out with --kv-cache-dtype bfloat16.
  • Cached blocks are not portable across the setting. A KV cache moved between instances (Disaggregation and KV Connectors) must be produced and consumed with the same cache type.

What does not change#

Quantization does not alter the scheduler, the block pool’s logic, or the prefix cache’s behaviour. Blocks are still 16 tokens; hashes are still over token IDs. Only the number of blocks and the time per step change. This is why it composes freely with every scheduling feature in this course.

Choosing#

SituationReasonable choice
Hopper/Ada or newer, quality-sensitiveFP8 W8A8 from LLM Compressor
Ampere (A100, A10G)INT4 W4A16 or AWQ through Marlin; or INT8 W8A8
Model does not fit at all4-bit weights, then measure quality on your own task
Model fits, cache is the constraintKeep weights as they are; --kv-cache-dtype fp8
No quantized checkpoint, want to tryOnline fp8_per_tensor
CPU servingINT8 W8A8

Always evaluate on your own workload. Aggregate benchmark scores hide task-specific losses, and 4-bit models in particular degrade unevenly: fluent prose survives, exact arithmetic and long structured outputs suffer first.

Code#

The serving effect in numbers: for one GPU and one model, how weight precision and cache precision change capacity and decode speed.

Go
package main

import "fmt"

const GiB = 1 << 30

func main() {
	const (
		gpuGiB       = 24.0
		util         = 0.92
		otherGiB     = 0.8  // activations, CUDA graphs, non-torch
		params       = 8.03 // billions of parameters
		kvPerTok16   = 131072.0
		avgReqTokens = 3000.0 // prompt + output of a typical request
		bandwidthGBs = 700.0  // GPU memory bandwidth, GB/s (order of magnitude)
	)

	type cfg struct {
		name       string
		weightBits float64
		kvBits     float64
		fastKernel bool // false: weights are unpacked to 16 bits on every use
	}
	cfgs := []cfg{
		{"16-bit weights, 16-bit KV", 16, 16, true},
		{"16-bit weights,  8-bit KV", 16, 8, true},
		{" 8-bit weights, 16-bit KV", 8, 16, true},
		{" 8-bit weights,  8-bit KV", 8, 8, true},
		{" 4-bit weights, 16-bit KV", 4, 16, true},
		{" 4-bit weights,  8-bit KV", 4, 8, true},
		{" 4-bit, no native kernel ", 4, 16, false},
	}

	fmt.Printf("%-28s %9s %9s %12s %12s %14s\n",
		"configuration", "weights", "KV pool", "tokens", "requests", "decode steps/s")
	for _, c := range cfgs {
		weights := params * 1e9 * c.weightBits / 8 / GiB
		kv := gpuGiB*util - weights - otherGiB
		tokens := kv * GiB / (kvPerTok16 * c.kvBits / 16)

		// A decode step must read every weight once: steps/s is bounded by bandwidth.
		readGB := params * c.weightBits / 8
		if !c.fastKernel {
			readGB = params*c.weightBits/8 + params*2 // read packed, then touch the 16-bit copy
		}
		fmt.Printf("%-28s %6.1f GiB %6.1f GiB %12.0f %12.0f %14.0f\n",
			c.name, weights, kv, tokens, tokens/avgReqTokens, bandwidthGBs/readGB)
	}
}

Three things to take from the table. Moving from 16-bit to 4-bit weights with a native kernel roughly quadruples the decode rate bound and nearly triples concurrent requests. Quantizing only the cache doubles concurrency without touching the weights. And the last row — 4-bit weights without a kernel that computes on them — has all the memory benefit and is slower than the 16-bit model.

Remember this#

  • Quantization helps a server twice: faster decode, and more memory left for KV blocks.
  • vLLM loads models quantized by other tools; the checkpoint’s config selects the scheme.
  • LLM Compressor is the project’s own tool and targets vLLM’s fast paths.
  • FP8 weights need Ada/Hopper-class GPUs or newer; 4-bit through Marlin runs on Ampere too.
  • A format without a native kernel saves memory but can be slower.
  • Online quantization converts a 16-bit model at load time with no calibration.
  • --kv-cache-dtype fp8 doubles cache capacity independently of the weights; calibrated scales improve its accuracy.
  • The scheduler and cache logic are unchanged; only block count and step time move.

Try it#

  1. Change the GPU to 80 GiB and the model to 70 billion parameters. Which configurations fit on one GPU at all?
  2. With avgReqTokens at 30,000 (long documents), which matters more for concurrency: weight precision or cache precision?
  3. On a real server, compare Available KV cache memory and tokens per second for the same model at 16 bits and with --quantization fp8_per_tensor.

Check yourself#

  1. Why does quantizing weights increase the number of requests a server can hold?
  2. An INT4 model is smaller but slower than the 16-bit original on your GPU. What is the likely reason?
  3. What does --kv-cache-dtype fp8 change, and what does it leave unchanged?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom