Preference Learning and Direct Preference Optimization
machine-learning
posttraining
Optimize relative preference between responses while measuring what the labels actually encode.
Demonstrations specify one acceptable response per prompt; preference pairs rank two. A preference example is (x, y_w, y_l): a prompt, a preferred response, and a rejected response. This chapter implements the numerical pieces behind a reward model and Direct Preference Optimization (DPO), then constructs a preference dataset whose labels deliberately contain a shortcut. The second evaluator stays untouched by optimization, so an improved training score can be compared with the behavior we actually intended.
The implementation uses small arrays rather than a language-model-sized network. That keeps every probability visible while preserving the same sequence-level objectives used by larger policies. Design rule: preference optimization is only as meaningful as the evaluator that remains outside the training loop.
Preference pairs
A reward model is trained to rank the preferred response above the rejected one:
DPO removes the separately trained reward model. With a trainable policy \pi_\theta and frozen reference \pi_{\mathrm{ref}}, it optimizes the difference in log-ratios,
The reference is not a second label source. It is a trust-region-like anchor that makes the update relative to a known policy. The temperature \beta controls how sharply a preference margin is pursued.
The sequence log-probability is the shared primitive: it sums token log-probabilities over the response mask, exactly as the SFT mask in Chapter 08 selected response positions.
Sequence log-probabilities
A sequence probability is a product of token probabilities, so its log-probability is a sum. Work in log space to avoid underflow and use the response mask to exclude prompt and padding tokens. The function below returns one scalar per sequence; changing a masked token cannot change its value.
The fourth position has different logits and targets in the two rows, but it is padding and contributes zero. The equality check is an invariant for batching: preference scores must not change when examples are padded to a common length. Summing rather than averaging is intentional here because DPO compares the likelihood of complete responses; if response lengths are allowed to differ, report a separate length analysis rather than silently changing the objective.
Pairwise reward models
A reward model maps a prompt-response sequence to a scalar. The tiny bag-of-tokens model below is not intended to understand language; it makes the pairwise update inspectable. Its embedding pools the response tokens and a linear head produces one reward. The training target is only the ordering, not an absolute score.
The reward model learns an ordering: all four preferred margins are positive on this tiny training set. That is a fit check, not evidence that the model learned helpfulness. The next experiment changes the data-generating process so the same optimization machinery can be shown to learn the wrong property.
Direct Preference Optimization (DPO)
For a full language model, the policy and reference log-probabilities come from two forward passes over the chosen and rejected responses. To isolate the DPO update, represent four complete responses by four trainable policy logits. The candidate log-probabilities below are the exact same quantities that the sequence log-probability function would return for a real decoder; only the model that produces them has been compressed to a lookup table.
DPO increases the policy’s relative likelihood of the preferred completions while the reference stays fixed. The loss does not say that the preferred response is factually correct; it says only that it should outrank the rejected response according to the labels. In practice, evaluate the policy at each checkpoint with held-out preference pairs and with a ground-truth evaluator that never supplies gradients.
Length bias
A preference dataset can make an accidental feature look like quality. Here every training winner is longer, so a reward model that sees only response length can achieve perfect training accuracy. The held-out pairs reverse the correlation: the preferred answer is concise. The evaluator is known in advance, which lets us measure the shortcut directly instead of debating whether a learned judge is trustworthy.
The model reaches perfect training accuracy by assigning a positive value to length, then fails every held-out pair because the intended preference is reversed. This is the smallest possible overoptimization example: the optimizer is behaving correctly with respect to the labels, while the labels fail to express the intended property. A real study should vary length independently of quality, report response-length distributions, and keep a ground-truth evaluator outside the reward or DPO loop.
Reliability checks
At every preference-training checkpoint, report:
held-out pairwise accuracy, separated by length-balanced and length-confounded subsets;
the ground-truth evaluator score, never used for gradients;
response length and token-level entropy;
the distance from the reference policy, such as a KL estimate;
SFT and general-suite scores from Chapters 07 and 08.
Noise, annotator disagreement, and reference choice should be swept explicitly. A falling DPO loss is compatible with a widening gap between reward-model score and the untouched evaluator. That gap is the measurement of proxy optimization, not an inconvenient outlier.
Shared DPO contract
The array implementation keeps the sequence objective visible; the reusable path takes chosen and rejected sequence log-probabilities relative to the same frozen reference. Because it consumes sequence scores rather than raw token positions, it can share the response-mask discipline from Chapter 08 and remain invariant to padding that is outside the response.
Pairwise reward modeling learns an ordering through -\log\sigma(r_w-r_l); it does not create an absolute notion of quality.
DPO updates relative policy likelihoods against a frozen reference, with \beta controlling the strength of the preference margin.
Sequence log-probabilities must use the same response mask discipline as SFT and must be invariant to padding.
A length-only evaluator can fit every training preference and fail every held-out preference. Always report a ground-truth evaluator, length-balanced subsets, and the general regression suite.
Chapter 10 turns the proxy problem into a policy-gradient experiment with an executable checker and a deliberate loophole.
Exercises
Use the exercises to test the chapter’s invariants and connect the derivations to the reusable implementation. Solutions are hidden in the notebook source and are available through the course tooling when needed.
[P9.1] Length confound
Suppose every preferred response in a dataset is exactly twice the length of its rejected counterpart. Describe what a reward model trained on this data learns to score, and specify one held-out pair that detects the confound.
[P9.2]
Why does DPO compare policy and reference log-ratios for both the chosen and rejected responses? State what the beta parameter changes.