Reinforcement Learning from Verifiable Rewards
Reinforcement Learning from Verifiable Rewards improves a generating policy using outcomes that a program can check, such as a correct mathematical answer or passing tests. The skill is designing reliable verifiers and a useful task distribution, then determining whether optimization teaches transferable problem solving or merely exploits the checks.
What it is
The model samples completions for tasks whose results can be evaluated without a learned human preference model. A verifier returns a reward based on properties such as answer equivalence, executable tests or valid output structure. Policy optimization increases the probability of rewarded completions; different reinforcement learning algorithms can perform this update. Outcome verification checks the final result and does not necessarily establish that an explanation is sound. Format rewards may supplement correctness rewards but represent a separate objective. The approach is most natural when evaluation is inexpensive, reproducible and difficult for the model to game, though these conditions need testing.
What the work involves
Build and version the verifier before scaling policy training. Inspect parsing, numerical tolerances, test coverage, timeouts and adversarial outputs that might receive credit incorrectly. Separate problem generators, templates and solutions across training and evaluation. Sample tasks at a difficulty that produces informative variation rather than uniformly zero or maximum reward. Track reward distributions and evaluate on independently checked problems. The deliverable includes a trained policy, a reproducible verification harness and evidence distinguishing actual task success from improvements caused by formatting or a weak checker.
Illustrative example
An illustrative coding exercise asks a model to implement a parser. The reward comes from unit tests in a sandbox. A developer adds cases for malformed input and verifies that a solution cannot read expected outputs from the harness. Training tasks and held-out parser specifications use different generators. The resulting policy is assessed on hidden tests and inspected for clear failure handling, because passing the training suite does not prove that it implements the specification.
Limits and common mistakes
Verification can be incomplete: tests miss behaviors, answer parsers accept ambiguous strings and numerical comparisons mishandle units. Sparse rewards make exploration difficult, while permissive checks invite shortcuts. Success on automatically checkable tasks does not establish quality on subjective assistance or open-ended factual claims. Correct answers also need not imply faithful reasoning traces. Treat verifier reliability and independent task accuracy as separate quality measures, and reassess them when the policy discovers new output strategies.
Prerequisites
- hardRLHF
RLVR swaps the human-preference reward for an automatic verifier.
Sits in the same preference/RL post-training family.
Sources and further reading
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Rule-based correctness and format rewards in reasoning-oriented reinforcement learning.
- Hugging Face TRL: GRPO Trainer
Custom reward functions, generated completion groups and policy training diagnostics.
Last updated: 2026-10-10