Skip to content
GenAI Learn/The Math You Actually Need
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

Broadcasting & Vector Norms

7 min read

You'll learn to

  • -Explain broadcasting as numpy's way of applying one operation to many rows at once
  • -Compute L1 and L2 norms and explain when each one matters
  • -Understand why embeddings get normalized before comparison

Real models never process one example at a time - they process a batch of thousands. Two ideas make that efficient: broadcasting, which lets one operation apply to every row of a batch at once, and norms, which measure a vector's size in ways that have real consequences for training and for comparing vectors to each other.

Broadcasting: One Line, Every Row

Suppose every example in a batch needs the same bias vector added to it. Looping over each row and adding the bias manually works, but numpy can do the entire batch in one call: shapes are compared starting from the right, so a bias of shape (d,) lines up automatically with the last dimension of a batch shaped (n, d), and gets "broadcast" across every row for free. The same trick, run the other way, computes every pairwise combination between two sets of vectors at once - the exact operation behind an all-pairs distance matrix, or a batch of attention scores.

The broadcasting idea, expressed with plain Python (numpy does this in one call, in compiled C)

L1 vs L2: Two Different Ideas of "Size"

A vector's "norm" is just a rule for measuring its size, and the rule you choose has real consequences. The L2 norm is the familiar straight-line length: square every element, sum, take the square root. The L1 norm sums the absolute values instead, with no squaring at all.

L1 and L2 Norms
‖x‖₁ = Σ|xᵢ| ‖x‖₂ = √(Σxᵢ²)
L1 sums absolute values. L2 is the familiar Euclidean length.
L1 norm, L2 norm, and normalizing a vector to unit length
Where Each Norm Actually Matters
Tends to push many weights to exactly zero - produces sparse models
L1 regularization (Lasso)
Shrinks weights smoothly toward zero without eliminating them
L2 regularization (Ridge)
Dividing by the L2 norm rescales a vector to unit length - the standard prep step before comparing embeddings by cosine similarity
Embedding normalization

AI Lab's Level 43 (Broadcasting a Batch) and Level 44 (L1 vs L2 Norms) implement both of these directly in numpy, including the exact reshape trick that makes an all-pairs comparison a single broadcasted expression.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement both from scratch: Level 43 (Broadcasting a Batch) and Level 44 (L1 vs L2 Norms) in AI Lab.