PPO, RLHF & How Reasoning Models Are Trained
You'll learn to
- -Explain why naive policy-gradient updates can be dangerously unstable for LLMs
- -Implement PPO's clipped surrogate objective
- -Understand the KL-penalized reward RLHF/GRPO/GSPO actually optimize
This is the chapter that ties the entire module together: the exact mechanisms behind RLHF and the reasoning-model training used by o1, DeepSeek-R1, and their successors. Two ideas, both building directly on what you just learned: PPO's clipped update, and a KL-divergence penalty that keeps a fine-tuned model from drifting into reward-hacked nonsense.
The Problem: An Unconstrained Update Can Wreck the Policy
Naively maximizing (new_policy_prob / old_policy_prob) * advantage sounds reasonable, but a single unusual batch can push that ratio far from 1 and take an update arbitrarily large in one step. For a small robotics policy that might just mean a bad episode. For a large language model, a wrecked update is not cheap to undo - the whole reason PPO (Proximal Policy Optimization) exists is to prevent exactly this.
PPO's Clipped Surrogate Objective
PPO caps how much the training objective can benefit from a probability ratio that has moved far from 1, by taking the minimum of the raw (unclipped) objective and a clipped version of it.
That last case is the subtle, interview-relevant part: the clip only removes the incentive to push a GOOD update (positive advantage) too far. If an action's advantage is negative, clipping does not protect the objective from getting even more negative - a genuinely bad update is never let off the hook.
RLHF and GRPO: Adding a KL Penalty
PPO alone still lets a policy drift arbitrarily far from where it started, chasing whatever a reward model rewards - which, left unconstrained, tends toward fluent-sounding, reward-hacked nonsense rather than genuinely better responses. RLHF fixes this by penalizing the reward itself with a KL-divergence term between the policy being trained and the original reference model, using the exact KL formula from earlier in this course.
AI Lab's Level 36 and 37 implement both formulas directly in numpy - and Level 37 reuses the exact KL divergence function you built all the way back in Level 25, now applied to a policy versus a reference model instead of two arbitrary distributions.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement both from scratch: Level 59 (PPO's Clipped Surrogate Objective) and Level 60 (KL-Penalized Reward) in AI Lab.