Serving: Making it fast and cheap
Between “we trained a great model” and “an API that serves a planet” sits some of the best systems engineering of our time — and it’s a shame how little of it most engineers see, because it’s the part of the stack built from ideas you already know: virtual memory, batch schedulers, caches, compression. This chapter reverse-engineers the price-per-million-tokens on the invoice.
The physics from last chapter, restated as the serving problem: decode is memory-bound — the GPU reads ~all active weights to emit each token, using ~1% of its compute. Every technique below is one of three moves: share the bytes across more work, shrink the bytes, or read the bytes less often.
Continuous batching: share the bytes
The waste in serving one user is that 140 GB of weights get read to produce one token. Serve 100 conversations in a single batch, and the same read produces 100 tokens — per-token cost falls ~100× before any cleverness. Throughput providers live and die on batch size (and this is why per-token API prices can sit so far below what a dedicated GPU would cost you).
The wrinkle is that requests don’t arrive in lockstep: one user’s answer is 10 tokens, another’s is 3,000. Naive batching waits for the longest member to finish — like an elevator that won’t open until every rider’s floor is reached. Continuous batching (also “in-flight batching”) schedules at the granularity of a single token step: finished requests exit the batch and queued ones enter, immediately. It’s a job scheduler where the quantum is one forward pass, and it’s table stakes in every serving engine.
The tradeoff it navigates is the classic one — throughput vs. latency. Bigger batches amortize better but make everyone wait a bit longer per token. The two numbers on every serving dashboard: TTFT (time to first token — dominated by prefill and queueing) and tokens/sec thereafter (dominated by decode). Batch APIs at half price are this tradeoff sold explicitly: your overnight jobs are the ballast that fills the batches.
PagedAttention: stop wasting the memory
Batching’s limiter is memory: every concurrent conversation carries its KV cache (Chapter 5), and the caches compete with the weights for HBM. Early servers allocated each request a contiguous max-length buffer — and delivered a museum-quality reproduction of pre-virtual-memory operating systems: internal fragmentation, external fragmentation, most of the “used” memory holding nothing.
The fix, from the vLLM paper , is called PagedAttention and it is exactly what its name promises: chop KV caches into fixed-size blocks (~16 tokens), allocate on demand from a shared pool, and translate through a per-request block table. Virtual memory, reinvented for attention — fragmentation vanishes, effective batch size jumps, and — just as page tables enabled shared memory — two requests with the same prefix can map the same physical blocks.
That last trick matters more than it sounds, because real traffic is wildly prefix-heavy: every request from your app shares the same long system prompt; every turn of a chat shares all previous turns; every step of an agent loop (next chapter) re-sends a growing transcript. Prefix caching — keeping hot prefixes’ KV blocks resident and skipping their prefill entirely — is why providers bill “cached input tokens” at a steep discount, and why you should structure prompts with the stable parts first (a genuinely actionable takeaway: put the variable parts of your prompt at the end).
At the frontier, serving splits further: prefill/decode disaggregation runs the compute-bound prefill and memory-bound decode on separate GPU pools, so neither workload’s utilization profile poisons the other’s.
Quantization: shrink the bytes
If decode speed is bytes-of-weights ÷ bandwidth, then halving the bytes doubles the ceiling. Quantization stores numbers in fewer bits — 16-bit floats down to 8-bit or 4-bit integers — by mapping each weight group’s range onto a small grid (plus per-group scale factors, and care for outlier values, which is where methods like GPTQ and AWQ earn their citations).
The empirical miracle: weights at 8-bit cost essentially nothing in quality, and 4-bit costs remarkably little — the giant undertrained matrices simply don’t need the precision (that a 70B model works at 4 bits/weight is itself a clue about how diffusely information is stored — Chapter 6 again). A 70B model at 4 bits is ~35 GB: one GPU instead of two, double the decode ceiling, or a frontier-class model on a MacBook — the entire local-LLM scene is downstream of this one technique. Quantizing the KV cache pays the same dividend to context length; quantizing activations (fp8) speeds prefill too. The costs are workload-dependent and show up in the tails, so the eval discipline of Chapter 9 applies to your quantized deployment as much as to any model swap.
Speculative decoding: read the bytes less often
The most conceptually delightful trick, from Leviathan et al. (and Chen et al. ): most tokens are easy — finishing “the United States of” doesn’t need a trillion parameters. But decoding makes the big model pay full price for every token.
So: let a small, fast draft model race ahead and propose several tokens, then have the big model check them all in one forward pass — which it can do because verifying a sequence is a parallel, prefill-like operation (that asymmetry between checking and generating is the whole trick). Accept the longest agreeing prefix; where the draft went wrong, the big model’s own prediction takes over. A careful accept/reject rule makes the output distribution provably identical to the big model’s — this is a lossless speedup, pure latency win, typically 2–3×. When the models agree on 4 tokens, you’ve paid one weight-read for four tokens: arithmetic intensity, raised again.
(Variants abound — extra prediction heads on the model itself, n-gram lookahead — and here’s the promised diffusion cameo: diffusion language models generate whole blocks of tokens in parallel and refine them iteratively, a bet on breaking the one-token-at-a-time bottleneck entirely rather than merely speculating past it. Watch that space; don’t reorganize around it.)
Distillation: shrink the model itself
The endgame move, older than everything above (Hinton et al., 2015 ): use a big teacher model to train a small student — on the teacher’s full output distributions (soft targets carry per-token information a hard label doesn’t) or, more commonly now, simply on millions of teacher-generated transcripts (Chapter 8’s SFT, with the teacher as the annotator). The mini/flash/lite tier of every model family is some blend of distillation and Chapter 4’s over-training; R1’s distilled small models are how reasoning reached laptops within weeks of its release.
Distillation is why “capability lag” keeps shrinking: whatever the frontier does expensively this quarter, a distilled model does cheaply next quarter. It’s also, notoriously, how competitors bootstrap from each other’s APIs — every lab’s terms of service forbid it, and every lab is suspected of it.
The stack, assembled
Trace a request through everything at once: it arrives, its system-prompt prefix hits cached KV blocks (no prefill), the novel suffix prefills on the compute pool, decode joins a continuously-batched swarm on the decode pool reading 4-bit weights, a draft model speculates 4 tokens ahead, and PagedAttention shuffles KV blocks underneath it all. None of these techniques knows anything about language. It’s caches, pages, schedulers, and compression — your field, aimed at one workload. The invoice line “per million tokens” is the price of bytes moved through HBM, divided by how cleverly they’re shared.
Further reading
- Efficient Memory Management for Large Language Model Serving with PagedAttention — the vLLM paper; the most satisfying systems read in the bibliography.
- Fast Inference from Transformers via Speculative Decoding — the rejection-sampling proof is short and elegant.
- GPTQ and AWQ — the canonical weight-quantization methods.
- Distilling the Knowledge in a Neural Network — Hinton’s original; a decade old and freshly relevant every year.