How Attention Actually Works
You'll learn to
- -Explain attention as similarity (dot product) followed by softmax
- -Implement scaled dot-product attention from scratch
- -Understand why real Transformers split attention into multiple heads
- -Explain why Grouped-Query Attention uses fewer Key/Value heads than Query heads, and what that saves
You have called an LLM through an API for most of this course, treating what happens inside it as someone else's problem. This chapter opens the box. Every Transformer layer, in every model from GPT-2 to the frontier model you used this morning, runs the exact same formula you are about to implement by hand.
The Question Attention Answers
For every position in a sequence, attention asks: "given what I'm looking for right now (my Query), how relevant is every other position (their Keys), and how should I blend their information (their Values) accordingly?" That is the entire idea. The rest is just the specific formula that makes it computable.
- -Q @ K.T scores every query against every key - the same dot-product similarity from earlier in this course, computed between every pair of positions at once.
- -Dividing by sqrt(d_k) keeps scores from growing huge as dimensionality increases, which would otherwise push softmax into an almost one-hot, hard-to-train regime.
- -softmax(scores) turns each query's row of scores into a probability distribution over every key.
- -Multiplying those weights by V blends every position's Value vector by how much attention the current position pays it - that blend IS the attention output.
Multi-Head Attention: Several Smaller Attentions in Parallel
A real Transformer layer does not run attention once - it runs it several times in parallel, in narrower slices called heads, so different heads can specialize (one might track syntax, another might track long-range topic coherence). The attention formula itself does not change at all. What is new is purely bookkeeping: split a wide vector into several narrow ones, run the same attention formula on each independently, then concatenate the results back together.
AI Lab's Level 26 and 27 implement exactly this in numpy: the vectorized attention formula across a whole sequence at once, then the reshape/transpose trick that splits and recombines attention heads.
Grouped-Query Attention: Fewer KV Heads, Cheaper to Serve
One more variant is worth knowing, because it is what most production models actually run: at generation time, a model caches every past token's Key and Value vectors for every head - the "KV cache" - and for a long conversation, that cache is often the single biggest memory cost of serving the model. Grouped-Query Attention (GQA) shrinks it directly: keep many Query heads for representational power, but far fewer Key/Value heads, and let groups of Query heads share each KV head.
- -Ordinary multi-head attention (what you just implemented) is the special case where every Query head has its own KV head - one-to-one.
- -Multi-Query Attention (MQA) is the opposite extreme: every Query head shares a single KV head.
- -GQA sits in between: num_groups = num_q_heads / num_kv_heads Query heads share each KV head - Llama 2 70B uses 64 Query heads but only 8 KV heads, an 8x smaller cache.
- -Once the KV heads are repeated to match the Query head count, the attention math is completely unchanged from Level 26/27 - GQA is a memory-and-serving optimization, not a different formula.
AI Lab's Level 39 builds this exact mechanism in numpy - repeating KV heads, then running the same batched attention formula across all Query heads at once.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement it in numpy: Level 49 (Scaled Dot-Product Attention), Level 50 (Multi-Head Attention's Reshape Trick), and Level 62 (Grouped-Query Attention) in AI Lab.