Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Preference learning: RLHF and DPO

By Anup Rai5 min readReviewed September 2026

Reinforcement learning from human feedback (RLHF) uses human feedback to guide a model through reinforcement learning. A common language-model recipe trains a reward model from human preferences and then optimizes a policy against that reward. Direct Preference Optimization (DPO) instead trains the policy directly on chosen/rejected response pairs using a preference objective relative to a reference policy.

Both can improve behavior on the feedback distribution. Neither is a proof of factuality, safety, or agreement with every human value. Begin an interview answer with the behavior being improved: appropriate abstention, instruction following, response quality, or another measurable goal.

1. Define the desired behavior

Consider a support assistant that confidently invents refund exceptions. The target is not “sound friendlier.” It is to answer supported questions accurately and abstain when the supplied policy is insufficient. Reviewers need a rubric that prefers grounded answers over fluent inventions.

  1. Specify which responses count as correct, unsupported, unsafe, or unhelpfully refused.
  2. Collect representative prompts, including ambiguous and adversarial cases.
  3. Obtain preference labels using that rubric; allow ties and disagreements.
  4. Keep a held-out evaluation that measures business behavior, not only a learned reward score.
  5. Compare against a prompting and supervised fine-tuning baseline.

2. Trace a common RLHF pipeline

Architecture / visual model
flowchart LR A[Instruction examples] --> B[Supervised policy] B --> C[Sample responses] C --> D[Human comparisons] D --> E[Train reward model] E --> F[RL policy optimization] F --> G[Independent evaluation]
Read diagram source
flowchart LR
  A[Instruction examples] --> B[Supervised policy]
  B --> C[Sample responses]
  C --> D[Human comparisons]
  D --> E[Train reward model]
  E --> F[RL policy optimization]
  F --> G[Independent evaluation]

The reward model learns which responses reviewers prefer. An algorithm such as PPO then updates the policy to increase expected reward, commonly with a penalty for drifting from a reference policy. PPO usually uses a learned value estimate; this adds training and memory costs. There are other RL algorithms and implementation choices, so “RLHF always needs exactly four full-size models” is too rigid. The InstructGPT paper documents one influential implementation.

A policy may learn to exploit reward-model errors. For example, a judge that rewards length may prefer a long, unsupported refund explanation. Track independent correctness, refusal quality, response length, and human evaluation alongside reward.

3. What DPO changes

For each prompt, DPO receives a preferred response and a rejected response. It increases the preferred response's relative likelihood compared with the rejected one, using a reference policy to define the comparison. The original formulation avoids training a separate explicit reward model and does not require generating fresh samples inside every offline optimization step. DPO paper.

Important distinction: increasing a preference margin does not require the absolute probability of every rejected response to decrease. The objective is a relative comparison. Its useful behavior depends on label quality, policy/reference initialization, the strength of the update, and coverage of the preference data.

Choice What it buys What it costs or misses
SFT on desired answers A direct target behavior to imitate Does not directly learn the chosen/rejected comparison
Offline DPO A comparatively simple training loop over preference pairs Static data may poorly represent the policy's current mistakes
Reward-model RL Optimizes responses sampled from the evolving policy More components, rollout cost, and reward-model exploitation risk
Online preference training Refreshes comparisons near current behavior Repeated generation, judging, labeling, and quality control

4. Online feedback and verifiable rewards

“Online” means training data or feedback is refreshed from the evolving policy. It does not mean a production chatbot should update weights after every customer message. Keep model versions, evaluation gates, and rollback decisions explicit.

RLOO is a policy-gradient approach with a leave-one-out baseline; it is not a synonym for online DPO. A judge model can provide preference labels, but its biases and errors must be measured against the intended rubric.

For code or mathematics, a verifier may score the outcome directly. RLVR can use passing tests or a checked answer as reward. This does not imply that reasoning models necessarily reward their intermediate reasoning rather than their final answer. Outcome supervision and process supervision are distinct choices.

5. Work through the failure and repair

Suppose 1,000 held-out support prompts produce these illustrative results:

Model Supported correct answers Unsupported answers Excessive refusals
Baseline 850 100 50
Candidate 870 30 100

The candidate reduces unsupported answers but doubles excessive refusals. Whether to launch depends on the risk and usefulness requirements, not a single preference score. Break down the changes by task, language, and policy ambiguity; inspect label disagreements; revise the training set; and repeat a protected evaluation.

An alignment tax is a reduction in some capabilities associated with behavioral adaptation. It is an observed tradeoff to measure, not an inevitable fixed penalty. A smaller update, broader training coverage, and better rubrics may help, but none replaces testing.

Interview practice

  1. How does DPO differ from SFT? SFT imitates target responses. DPO learns a preference comparison between chosen and rejected responses relative to a reference.
  2. Why not use the reward score as the launch criterion? The policy is trained to optimize it and may exploit it. Use independent measures of the desired behavior.
  3. What makes a useful preference pair? A clear difference under a stated rubric, relevant prompts, and consistent context—not merely one longer answer.
  4. Does DPO remove all alignment complexity? It simplifies the optimization pipeline; it does not fix poor feedback, distribution shift, or missing safety evaluation.
  5. When would online data help? When the improving policy produces failure modes poorly covered by the static pair set, provided the extra collection cost is justified.
  6. Do passing code tests prove faithful reasoning? No. Tests check selected observable outcomes. They may be incomplete and do not establish every intermediate claim.

Recall card and closing

Rubric → comparisons → objective → independent evaluation. State the signal being optimized, the deployment behavior being measured, and the cost of keeping them aligned. Keep authorization and guardrails in the application even after preference training.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← LoRA and QLoRA: understand what becomes smaller
NEXT LESSONKnowledge distillation: transfer useful behavior, measure the loss →

Explore the diagram