Skip to content
GenAI Learn/The Math You Actually Need
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

Entropy & KL Divergence

7 min read

You'll learn to

  • -Compute Shannon entropy and interpret it as a measure of uncertainty
  • -Compute KL divergence and explain why it is not symmetric
  • -Understand why RLHF penalizes a fine-tuned model with a KL-divergence term

Two quantities from information theory come up constantly once you look past basic loss functions: entropy, which measures how uncertain a probability distribution is, and KL divergence, which measures how far one distribution has drifted from another. The second one specifically shows up again the moment you look at how a fine-tuned language model is kept from drifting too far from where it started.

Entropy: How Spread Out Is a Distribution?

Entropy is maximized by a uniform distribution - maximum uncertainty, every outcome equally likely - and is zero for a distribution that puts all its probability on a single outcome, since there is nothing left to be uncertain about.

Shannon Entropy
H(p) = -Σ pᵢ log(pᵢ)
Sum of each outcome's probability times the (negative) log of that probability.
Entropy of a uniform distribution vs. a near-certain one

KL Divergence: How Far Has q Drifted From p?

KL divergence extends the same idea to compare two distributions instead of describing one. It is zero only when the two distributions are identical, and - this trips people up in interviews - it is not symmetric: KL(p || q) is generally not equal to KL(q || p).

KL Divergence
KL(p ‖ q) = Σ pᵢ log(pᵢ / qᵢ)
How much information is lost approximating p with q. Zero only when p and q are identical.
KL divergence between a fine-tuned policy and its original reference model

RLHF does not just maximize a reward model's score - left unconstrained, that reward-hacks into fluent-sounding nonsense. The fix subtracts beta * KL(policy || reference_model) from the reward being optimized: exactly the formula above, applied to a language model's own output distribution instead of two arbitrary ones. A larger beta keeps the fine-tuned model closer to its starting point; beta=0 removes the constraint entirely.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement both formulas yourself: Level 45 (Entropy & KL Divergence) in AI Lab.