Skip to content
GenAI Learn/Transformers Under the Hood
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

How Attention Actually Works

10 min read

You'll learn to

  • -Explain attention as similarity (dot product) followed by softmax
  • -Implement scaled dot-product attention from scratch
  • -Understand why real Transformers split attention into multiple heads
  • -Explain why Grouped-Query Attention uses fewer Key/Value heads than Query heads, and what that saves

You have called an LLM through an API for most of this course, treating what happens inside it as someone else's problem. This chapter opens the box. Every Transformer layer, in every model from GPT-2 to the frontier model you used this morning, runs the exact same formula you are about to implement by hand.

The Question Attention Answers

For every position in a sequence, attention asks: "given what I'm looking for right now (my Query), how relevant is every other position (their Keys), and how should I blend their information (their Values) accordingly?" That is the entire idea. The rest is just the specific formula that makes it computable.

Scaled Dot-Product Attention
Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V
Score every Query against every Key with a dot product, scale down, softmax into weights, then blend the Values by those weights.
  • -Q @ K.T scores every query against every key - the same dot-product similarity from earlier in this course, computed between every pair of positions at once.
  • -Dividing by sqrt(d_k) keeps scores from growing huge as dimensionality increases, which would otherwise push softmax into an almost one-hot, hard-to-train regime.
  • -softmax(scores) turns each query's row of scores into a probability distribution over every key.
  • -Multiplying those weights by V blends every position's Value vector by how much attention the current position pays it - that blend IS the attention output.
Scaled dot-product attention, one query at a time, in pure Python

Multi-Head Attention: Several Smaller Attentions in Parallel

A real Transformer layer does not run attention once - it runs it several times in parallel, in narrower slices called heads, so different heads can specialize (one might track syntax, another might track long-range topic coherence). The attention formula itself does not change at all. What is new is purely bookkeeping: split a wide vector into several narrow ones, run the same attention formula on each independently, then concatenate the results back together.

AI Lab's Level 26 and 27 implement exactly this in numpy: the vectorized attention formula across a whole sequence at once, then the reshape/transpose trick that splits and recombines attention heads.

Grouped-Query Attention: Fewer KV Heads, Cheaper to Serve

One more variant is worth knowing, because it is what most production models actually run: at generation time, a model caches every past token's Key and Value vectors for every head - the "KV cache" - and for a long conversation, that cache is often the single biggest memory cost of serving the model. Grouped-Query Attention (GQA) shrinks it directly: keep many Query heads for representational power, but far fewer Key/Value heads, and let groups of Query heads share each KV head.

  • -Ordinary multi-head attention (what you just implemented) is the special case where every Query head has its own KV head - one-to-one.
  • -Multi-Query Attention (MQA) is the opposite extreme: every Query head shares a single KV head.
  • -GQA sits in between: num_groups = num_q_heads / num_kv_heads Query heads share each KV head - Llama 2 70B uses 64 Query heads but only 8 KV heads, an 8x smaller cache.
  • -Once the KV heads are repeated to match the Query head count, the attention math is completely unchanged from Level 26/27 - GQA is a memory-and-serving optimization, not a different formula.
Fewer KV heads than Query heads - GQA's entire idea in one repeat

AI Lab's Level 39 builds this exact mechanism in numpy - repeating KV heads, then running the same batched attention formula across all Query heads at once.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement it in numpy: Level 49 (Scaled Dot-Product Attention), Level 50 (Multi-Head Attention's Reshape Trick), and Level 62 (Grouped-Query Attention) in AI Lab.