Similarity, Softmax & Attention Scores
You'll learn to
- -Compute cosine similarity and explain why it ignores magnitude
- -Implement softmax and explain why it amplifies the largest score
- -Connect both operations directly to how a Transformer's attention layer works
Two operations show up so often in modern AI that it is worth isolating them completely from everything else: measuring how similar two vectors are, and turning a list of raw scores into a probability distribution. Once you know these cold, a Transformer's attention mechanism stops looking mysterious - it is just these two operations, run at scale.
Cosine Similarity: Direction, Not Magnitude
You already met the dot product back in Vectors and What They Represent. Cosine similarity is the dot product with magnitude divided back out: dot(a, b) / (||a|| * ||b||). The result lives between -1 (opposite directions) and 1 (identical direction), and it answers a very specific, very useful question: ignoring how long two vectors are, how aligned is their direction? That is exactly the question a search engine asks when it ranks documents against a query, and exactly what a recommendation system asks when comparing user and item embeddings.
Softmax: Scores Into a Real Distribution
A model rarely produces probabilities directly - it produces raw, unbounded scores. Softmax is the standard way to turn any vector of scores into a proper probability distribution: exponentiate every score so nothing is negative, then divide by the total so everything sums to exactly 1. Because exponentials grow fast, a modest lead in raw score turns into a much larger lead in probability - softmax does not just rescale, it amplifies the winner.
Both Together: What Attention Actually Is
A Transformer's attention layer asks, for every position: "given what I'm looking for (my Query), how relevant is every other position (their Keys)?" It answers that with a dot product between Query and Key vectors - the same similarity measure from this chapter, just not normalized by magnitude the way cosine similarity is. Those raw scores then go through softmax, turning them into attention weights that sum to 1, which finally get used to blend everyone's Value vectors. Similarity plus softmax, applied at scale, is the entire mechanism.
If you want to build this exact mechanism with your own hands, AI Lab's Transformers From Scratch act (Levels 26-31) walks through scaled dot-product attention, multi-head attention, causal masking, and a full Transformer block, one piece at a time.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement both from scratch: Level 41 (Similarity with Dot Product) and Level 42 (Softmax Probabilities) in AI Lab.