Atlas · skill

RLHF

Reinforcement Learning from Human Feedback uses human judgments to shape a model's behavior through a learned reward signal. Practitioners design comparisons, train and test the reward model, and optimize a policy while controlling unwanted drift. The competence includes evaluating behavior independently of the score that training maximizes.

conceptAlignment

What it is

In a common language model pipeline, supervised demonstrations first establish useful response behavior. Annotators then compare candidate answers, and a reward model learns to assign scores that reflect those comparisons. A reinforcement learning algorithm updates the generating policy to obtain higher predicted reward, often with a penalty for divergence from a reference policy. PPO is one possible optimizer, rather than part of the definition. The broader method also applies to nonlanguage policies whose trajectories people evaluate. Human feedback expresses preferences under a particular rubric and sampling process; it does not directly reveal a universal or complete reward function.

What the work involves

Specify who supplies feedback and what tradeoffs the rubric asks them to make. Sample prompts and candidate answers that expose meaningful failures, preserve annotation disagreements and reserve comparisons for reward model validation. Monitor policy reward, divergence, response length and optimization stability, with a stopping rule based on independent evaluation. Hold out related conversations and tasks as groups. The useful result is a documented behavior change supported by fresh human comparisons and task checks, including cases where optimizing the learned reward harms another requirement.

Illustrative example

For an illustrative summarization assistant, reviewers prefer summaries that preserve an important caveat while omitting repetitive background. Their comparisons train a reward model. During policy optimization, the developer notices that highly scored summaries become longer, so a separate evaluation asks readers whether each summary covers the caveat within the requested length. The team selects a checkpoint using those judgments and source fidelity, rather than selecting the run with the highest reward.

Limits and common mistakes

A policy can exploit weaknesses in its reward model, and preference agreement on familiar examples may not transfer to new topics. Annotator population, instructions and candidate sampling all influence what is learned. Reward scores from different training runs are not automatically comparable. RLHF can improve measured preferences while leaving factual errors or competing objectives unresolved. Direct preference optimization and verifiable rewards are neighboring approaches with different learning signals, and should not be silently treated as the same pipeline.

Prerequisites

  • RLHF is applied AFTER SFT to align the model with preferences — SFT provides the baseline model that RLHF refines

  • RLHF uses PPO (a policy gradient RL algorithm) to optimize the language model against a reward model

Related skills

Sources and further reading

Last updated: 2026-10-10