Skip to content
GenAI Learn/Reinforcement Learning for LLMs
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

Policy Gradients & Advantage Estimation

8 min read

You'll learn to

  • -Explain REINFORCE's core idea: weight a log-probability gradient by the outcome it led to
  • -Understand why a baseline reduces variance without changing the expected gradient
  • -Compute an advantage and explain why it is preferred over a raw return

DQN learns Q-values and picks whichever action looks best. A policy gradient method takes a more direct path: it adjusts the probability of each action directly, based on how good the outcome was. This is the family of algorithms that everything from classic robotics RL to modern LLM reasoning training actually builds on.

REINFORCE: Reinforce the Good, Suppress the Bad

REINFORCE's idea, stated plainly: for an action that led to a high return, increase its log-probability; for one that led to a low return, decrease it - weighted by exactly how good or bad the outcome was, relative to some baseline.

REINFORCE Loss
L = -mean((Gₜ - b) · log π(aₜ))
Weight each action's log-probability by (return - baseline), then negate for gradient descent (minimizing this loss is the same as maximizing expected return).
The REINFORCE loss, from scratch

Why a Baseline? Variance, Not Bias

Subtracting a well-chosen baseline (often a learned value-function estimate) does not change what the gradient points toward on average - it is still an unbiased estimate of the true policy gradient. What it does change, dramatically, is the variance of that estimate. Without a baseline, a policy trained on noisy, high-magnitude returns updates erratically; with one, updates become far more stable. This is the entire reason baselines exist.

Advantage: How Much Better Than Expected?

The advantage function formalizes "baseline" into a specific, widely-used choice: the actual return minus a learned value function's prediction. It directly answers "how much better than expected was this outcome?" - exactly the signal a policy gradient method should train on.

Advantage
A(s, a) = G - V(s)
Actual return minus the value function's own prediction for that state.
Advantage, and standardizing a batch of advantages

Standardizing advantages across a batch (subtract the mean, divide by the standard deviation) is standard practice - it keeps the effective learning rate stable no matter the raw scale of the reward function being used.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement both from scratch: Level 57 (Policy Gradients / REINFORCE) and Level 58 (Advantage Estimation) in AI Lab.