Skip to content
GenAI Learn/Reinforcement Learning for LLMs
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

PPO, RLHF & How Reasoning Models Are Trained

9 min read

You'll learn to

  • -Explain why naive policy-gradient updates can be dangerously unstable for LLMs
  • -Implement PPO's clipped surrogate objective
  • -Understand the KL-penalized reward RLHF/GRPO/GSPO actually optimize

This is the chapter that ties the entire module together: the exact mechanisms behind RLHF and the reasoning-model training used by o1, DeepSeek-R1, and their successors. Two ideas, both building directly on what you just learned: PPO's clipped update, and a KL-divergence penalty that keeps a fine-tuned model from drifting into reward-hacked nonsense.

The Problem: An Unconstrained Update Can Wreck the Policy

Naively maximizing (new_policy_prob / old_policy_prob) * advantage sounds reasonable, but a single unusual batch can push that ratio far from 1 and take an update arbitrarily large in one step. For a small robotics policy that might just mean a bad episode. For a large language model, a wrecked update is not cheap to undo - the whole reason PPO (Proximal Policy Optimization) exists is to prevent exactly this.

PPO's Clipped Surrogate Objective

PPO caps how much the training objective can benefit from a probability ratio that has moved far from 1, by taking the minimum of the raw (unclipped) objective and a clipped version of it.

PPO Clipped Objective
L = mean(min(r·A, clip(r, 1-ε, 1+ε)·A))
r is the new/old policy probability ratio, A is the advantage. Taking the minimum removes the incentive to push r far from 1 when that would inflate the objective.
PPO's clipped objective, from scratch

That last case is the subtle, interview-relevant part: the clip only removes the incentive to push a GOOD update (positive advantage) too far. If an action's advantage is negative, clipping does not protect the objective from getting even more negative - a genuinely bad update is never let off the hook.

RLHF and GRPO: Adding a KL Penalty

PPO alone still lets a policy drift arbitrarily far from where it started, chasing whatever a reward model rewards - which, left unconstrained, tends toward fluent-sounding, reward-hacked nonsense rather than genuinely better responses. RLHF fixes this by penalizing the reward itself with a KL-divergence term between the policy being trained and the original reference model, using the exact KL formula from earlier in this course.

KL-Penalized Reward (RLHF)
R = reward_model_score - β · KL(π ‖ π_ref)
The reward actually optimized: the reward model's score, minus a KL-divergence penalty scaled by beta, between the current policy and the frozen reference model.
The RLHF-style reward, reusing the KL divergence formula from earlier in this course
What GRPO/GSPO Change
Computed from a GROUP of sampled responses to the same prompt, ranked against each other, instead of a separately trained value function
Advantage source
The same clipped-ratio objective and KL-penalty-against-a-reference-model mechanism from this chapter
What stays the same
This is the training approach behind DeepSeek-R1 and the broader wave of reasoning-focused models that followed it
Why it matters

AI Lab's Level 36 and 37 implement both formulas directly in numpy - and Level 37 reuses the exact KL divergence function you built all the way back in Level 25, now applied to a policy versus a reference model instead of two arbitrary distributions.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement both from scratch: Level 59 (PPO's Clipped Surrogate Objective) and Level 60 (KL-Penalized Reward) in AI Lab.