Atlas · skill

Reward Modeling

Reward modeling learns a scoring function from judgments about the quality of model outputs or actions. The competence is turning a defensible evaluation rubric into training data and a calibrated comparison model, then testing how reliably that proxy behaves on candidates outside the examples used to fit it.

conceptAlignment

What it is

A reward model takes a prompt and candidate response, or a state and trajectory, and produces a scalar or structured assessment. In pairwise language model setups, its parameters are trained so a preferred response receives a higher score than a rejected response. Other formulations use ratings, rankings or process-level judgments. Scores summarize learned preferences rather than objective correctness unless the data specifically supports that interpretation. A reward model may select candidates, evaluate experiments or supply a reinforcement learning objective. These uses expose it to different distributions, particularly when a policy actively searches for outputs that maximize its score.

What the work involves

Write a rubric, choose label granularity and gather comparisons with varied quality gaps. Check inter-annotator agreement and audit whether response length or formatting predicts the labels. Split related prompts and candidate sources together to avoid leakage. Evaluate pairwise accuracy and disagreement by task and response style; inspect score changes on controlled perturbations. Test generated candidates from policies that differ from the training generator. The result is a versioned scorer with a defined validity range and explicit evidence about its errors, suitable for a particular decision rather than an unrestricted quality oracle.

Illustrative example

An illustrative answer-selection system produces several explanations of a technical concept. Reviewers label pairs for factual accuracy and clarity. The fitted reward model chooses a candidate, but a challenge set adds polished answers containing a subtle false claim. The developer compares its rankings with fresh expert judgments and inspects whether the scorer values polish over correctness. Those findings determine whether it can automate selection or should only flag candidates for review.

Limits and common mistakes

High held-out comparison accuracy can hide systematic bias on rare but consequential cases. Scalar rewards collapse competing objectives, and their scale is usually meaningful only within the model's training setup. Optimization can amplify weaknesses that ordinary evaluation never encounters. A learned reward is distinct from a deterministic verifier and from direct policy preference losses. Validate under the intended selection or training pressure, preserve independent evaluators and avoid interpreting a high score as a probability that an answer is true.

Prerequisites

Sources and further reading

Last updated: 2026-10-10