Evaluation: Knowing if it worked
Every engineer knows you don’t ship code without tests. Now imagine your codebase is a trillion inscrutable floats (Chapter 6), its behavior changed in unknown ways by every training decision, and “correct” is defined as “answers arbitrary human requests well.” Evaluation is the testing discipline of machine learning, and for LLMs it is simultaneously the most important and most broken part of the field.
This chapter matters double for you, because unlike pre-training, you will do this. Anyone shipping an LLM feature is in the evals business, whether they’ve admitted it or not.
The benchmark canon
A benchmark is a dataset of tasks plus a grading scheme. A handful have defined the eras:
- MMLU  — ~16K multiple-choice questions across 57 subjects; the “general knowledge” headline number of the GPT-4 era.
- GSM8K  — grade-school math word problems; for years the standard reasoning test.
- HumanEval  — write a function from a docstring, graded by running the tests. Introduced pass@k (probability at least one of k samples passes), a genuinely better idiom: grade by execution, not text match.
- GPQA  — PhD-level science questions that are “Google-proof”; built explicitly because models saturated easier tests.
- SWE-bench  — resolve real GitHub issues in real repos, graded by the repo’s test suite. The closest canon benchmark to economically real work, and the industry’s favorite agent metric.
Read the progression: from multiple choice, to executed code, to do a junior engineer’s job. Benchmarks have had to chase capability upward — because models keep eating them.
Why you can’t trust the numbers
Four systematic rots, in ascending order of subtlety:
Contamination. Benchmarks are text on the internet; models train on the internet (Chapter 7’s decontamination is regex-vs-adversary and imperfect — a rephrased test question sails through). A model can “know” the test the way a student who found the answer key does. This is Chapter 1’s overfitting, resurrected at civilizational scale: the held-out set isn’t held out anymore. Partial defenses — private test sets, canary strings, continuously refreshed problems — all trade away reproducibility.
Saturation. Frontier models cluster in the high 80s and 90s on MMLU, where remaining gains are noise and mislabeled questions. A saturated benchmark ranks nothing; the canon has a shelf life, and the field must keep minting harder tests (GPQA exists precisely because MMLU died).
Goodhart’s law. The moment a benchmark appears in marketing, it stops measuring and starts steering. Labs tune data mixes with benchmarks in view; “benchmark-tuned” is a real slur for models that top leaderboards and disappoint users. The number was a proxy for capability; optimize the proxy and the gap between them widens — the exact reward-hacking dynamic of Chapter 8, one level up.
Construct validity. Does a 57-subject quiz measure what you mean by “smart”? Benchmark scores are precise measurements of something; whether that something is what you’re buying is a question the single headline number quietly skips.
Grading without answer keys: judges and arenas
Multiple choice and unit tests only cover tasks with checkable answers — Chapter 8’s “verifiable” end of the spectrum. For “write a good email,” grading needs judgment, and human judgment doesn’t scale. The field’s two answers:
LLM-as-judge. Use a strong model to grade outputs, pioneered by MT-Bench . It broadly agrees with human raters — with biases you must engineer around, and which should sound familiar from the reward-hacking discussion: position bias (favors the first answer shown; fix by grading both orders), length bias (favors longer), self-preference (favors its own family’s prose). A judge is a reward model you run at eval time — trust it accordingly, and never let a model family be its own final examiner.
Arenas. LMArena  (née Chatbot Arena) shows users two anonymous models’ answers and lets them vote, aggregating millions of votes into Elo ratings — the same math as chess. It resists contamination (prompts are live user traffic) and measures something real, but that something is “what impresses people in a chat window,” which conflates capability with confident formatting. Style-controlled variants exist precisely because the community learned that pleasing and correct are separable — Chapter 8’s sycophancy lesson, measured.
The eval stack you’ll actually build
Now the part that pays your rent. Public benchmarks tell you which model to start with; they say nothing about your task. Treat model choice and prompt changes the way you treat code changes — with a regression suite:
- Build a golden set. Fifty to a few hundred real examples from your product, each with a definition of “good.” Small and real beats large and synthetic; an afternoon of curation here is worth more than any leaderboard.
- Grade mechanically where possible. Exact match, schema validation, “did the code run,” “did it cite a real document” — cheap, deterministic checks catch a shocking fraction of regressions. Save judgment for what needs it.
- Use an LLM judge for the rest — and eval the judge. Write a rubric, grade both response orders, and spot-check judge verdicts against your own until you trust it (then keep spot-checking).
- Run it on every change. Model version bumps, prompt edits, retrieval tweaks — all of them shift behavior in ways that don’t announce themselves. The teams that ship reliable LLM features aren’t the ones with the best prompts; they’re the ones who notice when something breaks.
The mental shift: in ordinary software, tests verify logic you wrote. Here, the eval is the spec — it’s the only place where “what good looks like” is written down at all. Underinvest in it and you are, quite literally, developing without requirements.
A closing calibration
When a new model drops, the sophisticated read of its eval table: check the benchmarks you know aren’t saturated, discount anything the marketing leads with, treat arena Elo as a style-inclusive signal, and reserve judgment until it’s run on your golden set. Benchmarks are the field’s shared unit tests — imperfect, gamed, indispensable. The mistake isn’t trusting them too little or too much; it’s forgetting they were designed to answer someone else’s question.
Further reading
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena  — the judge paradigm, its biases, and the arena, in one paper.
- SWE-bench  — the benchmark that defined agentic coding evals; the grading design is worth studying for your own evals.
- GPQA  — a case study in building a contamination-resistant, expert-level benchmark.
- MMLU  and HumanEval  — the canon; read the grading sections, which shaped everything after.
- LMArena  — browse the live leaderboard with this chapter’s caveats in mind.