DPO & Test-Time Compute: Two More Ways to Get a Better Model
You'll learn to
- -Derive DPO's loss as a direct optimizer of the same KL-penalized objective RLHF/PPO optimize, without training a reward model
- -Explain why most open-model alignment moved from PPO to DPO
- -Implement Best-of-N and self-consistency as two simple test-time-compute techniques
- -Understand test-time compute as a separate axis from training compute for improving a reasoning model's answers
The last chapter built PPO and the KL-penalized reward RLHF optimizes - but PPO needs a separately-trained reward model and a rollout loop that is notoriously fiddly to tune. This chapter covers two techniques that sidestep that entirely, from two different directions: DPO skips PPO altogether during training, and test-time compute improves answers at inference time without touching training at all.
DPO: Optimizing the Same Objective, Without a Reward Model
DPO (Direct Preference Optimization) needs nothing but pairs of (chosen, rejected) human preference examples - no reward model, no PPO rollout loop. It works because the KL-penalized RLHF objective from the last chapter has a closed-form solution: given that objective, the optimal policy can be written directly in terms of the reference model and the reward, and rearranging that relationship turns "maximize reward under a KL constraint" into a simple, direct classification loss over preference pairs.
- -DPO needs four numbers per example: the policy's and the frozen reference model's log-probability of the chosen completion, and the same two for the rejected completion - never a reward-model score.
- -It is a binary classification loss in disguise: the policy is being trained to "classify" chosen as the preferred completion, exactly the way a logistic regression classifier is trained.
- -Because it needs no reward model, no sampling loop, and no PPO-style clipping, DPO trains with ordinary supervised-learning-style gradient descent - which is why most open-model alignment (Llama, Zephyr, and others) moved to DPO instead of PPO for most of their alignment work.
DPO and PPO/RLHF are not rivals with different goals - they are two different ways of reaching the same KL-constrained objective. PPO reaches it iteratively through sampling and a reward model; DPO reaches it in one closed-form step directly from preference pairs.
Test-Time Compute: Getting Smarter Without Retraining
Everything so far in this module improves a model by spending more compute on TRAINING it. Reasoning models like o1 and DeepSeek-R1 also spend extra compute at INFERENCE time: sampling multiple candidate solutions to the same problem, then aggregating over them, instead of trusting a single greedy pass. Two of the simplest aggregation strategies:
- -Best-of-N: sample N completions, score each with a reward model or verifier, and keep the highest-scoring one.
- -Self-consistency: sample N independent reasoning chains, extract a final answer from each, and return whichever answer appears most often - on the theory that correct reasoning tends to converge on the same answer while flawed reasoning scatters.
Both techniques trade inference cost (N samples instead of 1) for accuracy - this is exactly why reasoning models are slower and more expensive per query than a standard chat model, and why "test-time compute scaling" is now discussed as its own lever, alongside training compute and model size.
AI Lab's Level 43 and 44 implement both of this chapter's ideas in numpy - the DPO loss over a batch, and vectorized Best-of-N / self-consistency using np.argmax and np.unique.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement both from scratch: Level 66 (Direct Preference Optimization) and Level 67 (Best-of-N & Self-Consistency) in AI Lab.