Policy Gradients & Advantage Estimation
You'll learn to
- -Explain REINFORCE's core idea: weight a log-probability gradient by the outcome it led to
- -Understand why a baseline reduces variance without changing the expected gradient
- -Compute an advantage and explain why it is preferred over a raw return
DQN learns Q-values and picks whichever action looks best. A policy gradient method takes a more direct path: it adjusts the probability of each action directly, based on how good the outcome was. This is the family of algorithms that everything from classic robotics RL to modern LLM reasoning training actually builds on.
REINFORCE: Reinforce the Good, Suppress the Bad
REINFORCE's idea, stated plainly: for an action that led to a high return, increase its log-probability; for one that led to a low return, decrease it - weighted by exactly how good or bad the outcome was, relative to some baseline.
Why a Baseline? Variance, Not Bias
Subtracting a well-chosen baseline (often a learned value-function estimate) does not change what the gradient points toward on average - it is still an unbiased estimate of the true policy gradient. What it does change, dramatically, is the variance of that estimate. Without a baseline, a policy trained on noisy, high-magnitude returns updates erratically; with one, updates become far more stable. This is the entire reason baselines exist.
Advantage: How Much Better Than Expected?
The advantage function formalizes "baseline" into a specific, widely-used choice: the actual return minus a learned value function's prediction. It directly answers "how much better than expected was this outcome?" - exactly the signal a policy gradient method should train on.
Standardizing advantages across a batch (subtract the mean, divide by the standard deviation) is standard practice - it keeps the effective learning rate stable no matter the raw scale of the reward function being used.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement both from scratch: Level 57 (Policy Gradients / REINFORCE) and Level 58 (Advantage Estimation) in AI Lab.