Skip to Content
Chapter 3 · 7 minute read

The transformer: Attention is all you need

In 2017, a team at Google working on machine translation published a paper with an unusually cocky title: Attention Is All You Need . The claim was that the attention mechanism — until then a bolt-on used to help RNNs translate — could replace recurrence entirely. The resulting architecture, the transformer, is the substance of every model in this book. GPT stands for Generative Pre-trained Transformer.

Nine years later, the architecture is nearly unchanged. Understand this chapter and you understand the object everything else in the book trains, scales, interprets, and serves.

The problem attention solves

Recall the RNN’s flaws: it crushed the past into one fixed-size vector, and it processed tokens serially. The transformer’s move is to keep every token’s representation around, and let each token directly look at all the ones before it. No summarizing bottleneck, no recurrence — and because each token’s computation no longer depends on the previous token’s output, every position in the sequence can be processed in parallel during training. That second property is arguably the more important one: it’s what let the architecture soak up all the compute the next chapter is about.

Attention: a differentiable hash map

Here’s the mechanism, in the vocabulary of a systems engineer.

Each token’s current vector produces three new vectors via learned linear projections:

  • a query qq — “what am I looking for?”
  • a key kk — “what can I be found by?”
  • a value vv — “what information do I carry?”

To update a given token, take its query and dot-product it against the key of every preceding token. A big dot product means “this token is relevant to me.” Softmax those scores into weights that sum to 1, then take the weighted average of the corresponding values. That average gets added into the token’s representation:

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \tip{turn relevance scores into weights that sum to 1}{\mathrm{softmax}}\!\left(\frac{\tip{every query dotted with every key: how relevant is each token to each other token}{QK^\top}}{\tip{keeps scores from growing with vector size and saturating the softmax}{\sqrt{d_k}}}\right) \tip{the weighted average of the values — what actually gets retrieved}{V}

(The dk\sqrt{d_k} just keeps dot products from growing with vector dimension and saturating the softmax.)

The analogy I find most useful: attention is a hash map with soft lookups. A normal map retrieves exactly one value for a key. Attention retrieves every value at once, weighted by how well the query matches each key — and because the whole thing is differentiable, gradient descent can learn what to ask for, what to advertise, and what to carry. When the model processes it in “the cat sat on the mat because it was tired,” some attention head is emitting a query like “what noun am I about?” that matches the key advertised by cat, and pulls in cat’s information.

the0.03cat0.58sat0.06on0.02the0.02mat0.08because0.05it →was0.05tired0.11
One attention head resolving a pronoun: the query from "it" matches the key advertised by "cat", so cat's value dominates the weighted average. Shading shows the attention weights.

Because queries, keys, and values are all just matrices, the entire operation is a handful of matrix multiplies:

def attention(x, W_q, W_k, W_v): # x: (seq_len, d_model) Q, K, V = x @ W_q, x @ W_k, x @ W_v scores = Q @ K.T / math.sqrt(K.shape[-1]) # (seq_len, seq_len) scores = causal_mask(scores) # -inf above the diagonal return softmax(scores, axis=-1) @ V

Two details in that snippet deserve names:

  • Causal masking: during training, all positions are computed at once, so position 5 could “cheat” by attending to position 6. Setting the upper triangle of the score matrix to −∞-\infty before the softmax prevents this — every token only sees its past. This is what makes the model autoregressive.
  • The quadratic: that (seq_len, seq_len) score matrix means attention costs grow with the square of context length. A huge amount of the engineering in Chapters 5, 10, and 11 exists to manage this.

Multi-head attention is the observation that one soft lookup per token is not enough: you want one head tracking syntax, another tracking coreference, another matching brackets in code. So the model runs many attention “heads” in parallel (each with its own smaller Q/K/V projections) and concatenates the results. A frontier model might have ~100 heads per layer.

Position: attention is a set operation

Look at the code again: nothing in it knows where a token is. Shuffle the keys and values, and each output is unchanged — attention treats context as a set, not a sequence. Word order has to be injected explicitly.

The original paper added sinusoidal “positional encodings” to the embeddings. Modern models instead use rotary position embeddings (RoPE ), which rotate each query and key vector by an angle proportional to its position — so the dot product between a query and key ends up depending on their relative distance, which is what you actually want (“the token 3 back” matters; “token #4,821” doesn’t). We’ll return to RoPE when we talk about long context in Chapter 5.

The MLP returns

Attention moves information between tokens, but does surprisingly little computation on it. After each attention step, every token is passed — independently, in parallel — through an MLP: the exact architecture from Chapter 1, typically expanding to 4× the model dimension and back.

The division of labor is worth engraving: attention routes, the MLP computes. Attention gathers relevant context into a token’s vector; the MLP is where that vector gets processed — and, as we’ll see in Chapter 6, where most of the model’s factual knowledge appears to be stored. In a standard transformer the MLPs hold roughly two-thirds of the parameters. When people say “a bigger model knows more,” they are mostly talking about MLP weights.

The residual stream

One attention step plus one MLP (each preceded by a normalization step that keeps vector magnitudes stable) makes a block. A model is a stack of these blocks — from 12 in GPT-2 to on the order of a hundred in frontier models.

Crucially, each block doesn’t replace the token’s vector; it computes a modification and adds it:

x←x+Attention(x);x←x+MLP(x)\tip{the token's residual stream}{x} \leftarrow x + \tip{reads the stream and writes back what it gathered from other tokens}{\mathrm{Attention}(x)}; \qquad x \leftarrow x + \tip{reads the stream and writes back what it computed on this token}{\mathrm{MLP}(x)}

These additive shortcuts (“residual connections”) were invented to keep gradients flowing through deep networks, but they enable a mental model that the interpretability field (Chapter 6) has made standard: each token has a residual stream — a vector flowing up through the layers, acting like a shared memory bus. Every attention head and MLP reads from the stream, computes something, and writes its result back by addition. Early layers write low-level features, middle layers write concepts and facts, late layers convert it all into a concrete next-token prediction.

token embedding (enters)residual streamattention+MLP+moves info between tokenscomputes on each token× N layers
One transformer block, drawn the interpretability way: the residual stream flows upward like a memory bus; attention and the MLP each read from it, compute, and add their result back. A model is this, stacked N times.

So here is the entire architecture, end to end:

  1. Look up each token’s embedding (Chapter 2).
  2. Pass the sequence through N blocks of attention → MLP, each reading from and writing to the residual stream.
  3. Project the final vector of the last token onto the vocabulary to get logits, softmax into a distribution (Chapters 1–2), and predict.

That’s it. There is no module labeled “grammar,” no knowledge base, no planner. Everything a frontier model does is implemented in the weights of these identical, repeated blocks — which is either deeply unsatisfying or the most interesting fact in computer science, depending on your mood.

Encoders, decoders, and the lineage that won

The 2017 paper had two towers (an encoder for the source language, a decoder for the target). The years since sorted their descendants: BERT -style encoders powered search and classification; OpenAI’s GPT line took the decoder-only path — just the causal stack described above, trained on next-token prediction. GPT-2  (2019) showed one model could do many tasks without task-specific training; GPT-3  (2020) showed what happens when you scale it 100×. Every frontier LLM today is decoder-only. Why scaling was the winning bet is the next chapter.

One practical seed to plant before we go: notice that when generating token 1,001, the keys and values of tokens 1–1,000 haven’t changed — recomputing them would be pure waste. Caching them (the KV cache) is the single most important fact of inference economics, and Chapter 11 is largely about it.

Further reading