Post-training: From autocomplete to assistant
GPT-3 existed for two and a half years before ChatGPT. The model that took over the world was not much more capable than the base model that preceded it — it was more usable, and the difference was post-training. If pre-training pours in the knowledge, post-training shapes the thing you actually talk to: the assistant persona, the instruction-following, the refusals, and — since 2024 — the reasoning.
A useful frame before diving in: the base model is a simulator of documents, and the assistant is a character the model has been trained to play. Post-training is character design, wielding less than 1% of the total compute — a lever whose smallness should surprise you every time you think about it.
Supervised fine-tuning: show it the format
The first step is embarrassingly direct. Hire people (increasingly: use stronger models, with humans editing) to write ideal transcripts — conversations between a user and the assistant you wish existed. Then train on them with the exact same next-token loss from Chapter 1.
This is supervised fine-tuning (SFT), and mechanically there is nothing new in it: it’s more pre-training, on a tiny, curated corpus of “documents” that all happen to depict a helpful assistant. (If you learned some ML years ago, you may know this whole pattern — train a general network once, expensively, then adapt it to tasks cheaply — as transfer learning. In the LLM era nobody says the term anymore, because it stopped being a technique and became the field’s entire structure: the “P” in GPT stands for pre-trained precisely because the model is built to be adapted.) Conversations are serialized with special tokens marking who’s speaking — the chat template:
<|system|>You are a helpful assistant.<|end|>
<|user|>What's the capital of France?<|end|>
<|assistant|>Paris.<|end|>(When you call a chat API, this is what your list of messages compiles to. The model is still just completing a document — the document is a screenplay, and the API stops generation when the model emits the end-of-turn token.)
A few thousand to a few million high-quality examples teach the model the format and persona of assistance. What SFT can’t teach well is judgment — the difference between a good answer and a great one — because writing perfect demonstrations of judgment is exactly as hard as having it.
RLHF: grading is easier than writing
The insight that unlocked ChatGPT-level behavior, from the InstructGPT  paper: people are much better at comparing answers than producing them. You may not be able to write the perfect explanation of a tax form, but shown two candidates, you can reliably say which is better. Reinforcement learning from human feedback (RLHF) converts that cheap signal into training pressure, in two steps:
- Train a reward model. Sample pairs of model outputs, have humans pick the better one, and train a separate model (typically a clone of the LLM with the token-prediction head swapped for a single score output) to predict human preference. You’ve now compressed human judgment into a function you can call millions of times.
- Optimize against it with RL. The LLM generates responses, the reward model scores them, and an algorithm — classically PPO — adjusts the LLM to make high-scoring responses more probable. Unlike SFT’s “reproduce this exact text,” RL says “produce whatever scores well” — the model can discover behaviors no human demonstrated. (This is reinforcement learning’s original lineage: the same family of ideas from learning Atari and robot control from human preferences , repurposed.)
One guardrail matters enough to name: a KL penalty tethers the updated model to its pre-RL self, punishing it for drifting too far. Without the leash, the model finds the reward model’s bugs — because the reward model is a lossy proxy for what humans want, and a sufficiently strong optimizer will exploit the gap. This failure mode, reward hacking, is post-training’s central demon: optimize “answers humans rate highly” hard enough and you get confident-sounding, agreeable, flattering answers — the well-documented sycophancy of chat models is widely attributed to exactly this. Goodhart’s law, running at a million samples an hour.
Two important variations:
- Direct Preference Optimization  (DPO) collapses the pipeline: a clever loss lets you train on preference pairs directly, no reward model, no RL loop. Far simpler and cheaper; the workhorse of open-source post-training. Frontier labs still mostly run the full RL pipeline — with online sampling and continually retrained reward models — because it can iterate past the fixed dataset that DPO is anchored to.
- Constitutional AI  (Anthropic) substitutes AI feedback for most human feedback: a written constitution of principles, a model that critiques and revises outputs against it, and preference labels generated by the model itself (RLAIF). The scalable-oversight bet: use models to supervise models, keeping humans at the top writing principles instead of grading millions of transcripts.
RLVR and reasoning: the 2024 turn
Everything above optimizes for human approval — a fuzzy, hackable signal. The breakthrough behind reasoning models was to find domains where the reward can be verified: the math answer is right or wrong, the code passes the tests or doesn’t. Reinforcement learning from verifiable rewards (RLVR) is RLHF with the squishy reward model replaced by a checker that cannot be flattered.
Given an unhackable reward and room to explore, something remarkable happens — documented most openly in DeepSeek’s R1 paper : the model discovers extended reasoning on its own. Trained only on “was the final answer right,” R1’s precursor spontaneously learned to produce long chains of thought, to double-check itself, to backtrack — visible as an “aha-moment” in training where response lengths and accuracy climb together. Nobody demonstrated reasoning transcripts to it; thinking longer simply wins under a verifiable reward, so the optimizer found it. (OpenAI’s o1  announced the same regime; R1 showed everyone how.)
This is the third scaling axis promised in Chapter 4 — test-time compute — made concrete: reasoning models convert inference tokens into accuracy, and the exchange rate is good enough that “think for 10,000 tokens” competes with “be 10× bigger.” It also reframes the economics of Part IV: a reasoning model’s cost lives in its inference bill, not just its training bill.
Two caveats to keep your model of reasoning models honest. Verifiability is a spectrum — math and code sit at the easy end, and extending RLVR to essay quality or medical advice requires reward models again (with all their hackability; there’s active work on making the graders themselves reason). And recall Chapter 6: the chain of thought is evidence of the computation, not a faithful log of it.
The practical tier: fine-tuning for the rest of us
You will probably never run RLHF, but you may well fine-tune. LoRA  (low-rank adaptation) makes it economical: freeze the model, and train only small low-rank matrices added to each weight — under 1% of the parameters, a single GPU instead of a cluster, and swappable adapters per customer or task at serving time.
The advice half of this section: fine-tuning teaches form extremely well (your JSON schema, your tone, your domain’s jargon) and facts poorly — new knowledge injected via fine-tuning tends to be brittle and to collide with what the model already believes. For facts, retrieval (Chapter 12) usually beats weights. Fine-tune for behavior; retrieve for knowledge. And know the tax you’re paying either way: any narrow training erodes some general capability (catastrophic forgetting), which is why labs mix general data into every specialization run — and why a model that’s been aggressively aligned can feel slightly dumber than its base (the “alignment tax,” real but much reduced in modern pipelines).
What post-training can’t fix
A clean mental model to close the part on: pre-training sets the ceiling — the knowledge and raw capability. Post-training determines how much of the ceiling you can reach through the chat interface, and RLVR genuinely raises it for verifiable domains. But no amount of preference tuning installs knowledge that isn’t there; hallucination in particular is managed by post-training (calibrating “I don’t know”), never eliminated by it. When a model fails you, it’s worth asking which layer failed: the ceiling, or the character.
Further reading
- Training language models to follow instructions with human feedback  — InstructGPT: the paper that turned GPT-3 into a product; the RLHF pipeline in full.
- Direct Preference Optimization  — the simplification that democratized preference tuning.
- Constitutional AI  — RLAIF and the scalable-oversight bet.
- DeepSeek-R1  — the most open account of RLVR and emergent reasoning; read the “aha moment” section.
- RLHF: Reinforcement Learning from Human Feedback  — Chip Huyen’s explainer; the best single narrative of the classical pipeline.
- LoRA  — the fine-tuning method you’d actually use.