Skip to content
GenAI Learn/Reinforcement Learning for LLMs
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

DPO & Test-Time Compute: Two More Ways to Get a Better Model

9 min read

You'll learn to

  • -Derive DPO's loss as a direct optimizer of the same KL-penalized objective RLHF/PPO optimize, without training a reward model
  • -Explain why most open-model alignment moved from PPO to DPO
  • -Implement Best-of-N and self-consistency as two simple test-time-compute techniques
  • -Understand test-time compute as a separate axis from training compute for improving a reasoning model's answers

The last chapter built PPO and the KL-penalized reward RLHF optimizes - but PPO needs a separately-trained reward model and a rollout loop that is notoriously fiddly to tune. This chapter covers two techniques that sidestep that entirely, from two different directions: DPO skips PPO altogether during training, and test-time compute improves answers at inference time without touching training at all.

DPO: Optimizing the Same Objective, Without a Reward Model

DPO (Direct Preference Optimization) needs nothing but pairs of (chosen, rejected) human preference examples - no reward model, no PPO rollout loop. It works because the KL-penalized RLHF objective from the last chapter has a closed-form solution: given that objective, the optimal policy can be written directly in terms of the reference model and the reward, and rearranging that relationship turns "maximize reward under a KL constraint" into a simple, direct classification loss over preference pairs.

DPO Loss
L = -log σ(β · [(log πθ(yw|x) - log πref(yw|x)) - (log πθ(yl|x) - log πref(yl|x))])
yw = chosen ("winning") completion, yl = rejected ("losing") completion. The policy is rewarded for preferring chosen over rejected MORE than the reference model already did, scaled by beta - the exact same beta that controlled the KL penalty in PPO/RLHF.
The DPO loss, from scratch
  • -DPO needs four numbers per example: the policy's and the frozen reference model's log-probability of the chosen completion, and the same two for the rejected completion - never a reward-model score.
  • -It is a binary classification loss in disguise: the policy is being trained to "classify" chosen as the preferred completion, exactly the way a logistic regression classifier is trained.
  • -Because it needs no reward model, no sampling loop, and no PPO-style clipping, DPO trains with ordinary supervised-learning-style gradient descent - which is why most open-model alignment (Llama, Zephyr, and others) moved to DPO instead of PPO for most of their alignment work.

DPO and PPO/RLHF are not rivals with different goals - they are two different ways of reaching the same KL-constrained objective. PPO reaches it iteratively through sampling and a reward model; DPO reaches it in one closed-form step directly from preference pairs.

Test-Time Compute: Getting Smarter Without Retraining

Everything so far in this module improves a model by spending more compute on TRAINING it. Reasoning models like o1 and DeepSeek-R1 also spend extra compute at INFERENCE time: sampling multiple candidate solutions to the same problem, then aggregating over them, instead of trusting a single greedy pass. Two of the simplest aggregation strategies:

  • -Best-of-N: sample N completions, score each with a reward model or verifier, and keep the highest-scoring one.
  • -Self-consistency: sample N independent reasoning chains, extract a final answer from each, and return whichever answer appears most often - on the theory that correct reasoning tends to converge on the same answer while flawed reasoning scatters.
Best-of-N and self-consistency, from scratch

Both techniques trade inference cost (N samples instead of 1) for accuracy - this is exactly why reasoning models are slower and more expensive per query than a standard chat model, and why "test-time compute scaling" is now discussed as its own lever, alongside training compute and model size.

AI Lab's Level 43 and 44 implement both of this chapter's ideas in numpy - the DPO loss over a batch, and vectorized Best-of-N / self-consistency using np.argmax and np.unique.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement both from scratch: Level 66 (Direct Preference Optimization) and Level 67 (Best-of-N & Self-Consistency) in AI Lab.