Causal Masking & Positional Encoding
You'll learn to
- -Explain why an LLM must not attend to future tokens during generation
- -Implement a causal mask using broadcasting-style comparisons
- -Understand why positional encoding is needed at all, and how the sin/cos pattern works
- -Implement RoPE and explain why rotating Q/K beats adding a fixed pattern to the embedding
The attention mechanism from the last chapter has two problems left to solve before it can generate text one token at a time. First: nothing stops a position from attending to a position that comes after it, which would let the model cheat during training by peeking at the answer. Second: attention treats its input as an unordered set - swap two tokens and, left alone, attention produces the same set of outputs just reordered. Word order clearly matters. Both problems have small, elegant fixes.
Causal Masking: No Looking Ahead
When an LLM generates text, token 5 must never be allowed to attend to token 8 - it has not been generated yet. The fix is a mask added directly to the raw attention scores, before softmax: 0 (unchanged) for any position on or before the current one, negative infinity for anything after it.
The reason -infinity specifically: exp(-infinity) is exactly 0, so once the mask is added to the scores and softmax runs, every blocked position gets exactly zero attention weight. No separate "zero it out afterward" step is needed - the mask does its job before softmax even runs.
Positional Encoding: Injecting Word Order
Since attention alone is order-blind, the original Transformer paper injects position directly into the input: a fixed, unique pattern of sines and cosines added to every token's embedding, one pattern per position.
- -Position 0 always encodes to [0, 1, 0, 1, ...] regardless of the model's dimensionality, since sin(0) = 0 and cos(0) = 1 everywhere.
- -Low dimensions oscillate quickly - useful for distinguishing nearby positions.
- -High dimensions oscillate slowly - useful for distinguishing far-apart positions.
- -Together, the full pattern lets the model infer both exact and relative position from the embedding alone.
RoPE: What Llama, Mistral, and DeepSeek Actually Use
Sinusoidal encoding is how the original Transformer paper injected position - but it is not what any major open-weight model uses today. Rotary Positional Embeddings (RoPE) keeps the exact same frequency schedule (theta_i = base^(-2i/d)) but changes the OPERATION: instead of adding a fixed pattern to the embedding, RoPE ROTATES pairs of embedding dimensions by an angle proportional to position.
- -Position 0 always rotates by angle 0 (cos 0 = 1, sin 0 = 0), leaving the vector unchanged - the same sanity check the sinusoidal version had, for the same reason.
- -The payoff for switching from addition to rotation: the dot product between a rotated Query and a rotated Key ends up depending only on their RELATIVE distance (their position difference), not their absolute positions.
- -That relative-distance property is exactly what attention benefits from, and an additive encoding does not give you as cleanly - it is the specific reason RoPE displaced sinusoidal PE across the field.
AI Lab's Level 38 implements this across a full (seq_len, d_model) array at once, using the same broadcasting pattern as Level 30's sinusoidal encoding - only the final rotation step differs.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement all three from scratch: Level 51 (Causal Masking), Level 53 (Sinusoidal Positional Encoding), and Level 61 (RoPE) in AI Lab.