GRPO
Group Relative Policy Optimization updates a generating policy using rewards compared within groups of sampled responses. The skill is constructing informative prompt groups, implementing reliable rewards and controlling policy changes. GRPO can support verifiable or learned rewards; its optimizer is distinct from the source of the reward signal.
What it is
GRPO samples multiple completions for a prompt and uses their rewards to estimate relative advantages, commonly by centering and scaling scores within the group. It avoids a separately trained value model in the original formulation, using the group as its baseline instead. A clipped policy objective limits updates relative to the sampling policy, and formulations may include a penalty for divergence from a reference. Group size, generation settings and reward variability affect the learning signal. Modern implementations offer normalization and loss variants, so the exact objective needs documentation. GRPO does not define whether correctness, human preference or another property supplies the reward.
What the work involves
Test reward functions independently and select prompts where sampled completions exhibit meaningful quality differences. Choose group size and generation limits within the available inference and training budget. Inspect zero-variance groups, reward distributions, completion lengths, clipping and divergence. Keep problem families and solution sources separate from evaluation. Record the implementation's loss and normalization settings, then assess independently verified task accuracy rather than average reward alone. The useful result is a policy improvement with a reproducible sampling and optimization recipe and evidence that the relative signal corresponds to the intended behavior.
Illustrative example
An illustrative arithmetic policy generates several solutions for each exercise. A verifier rewards final answers only when parsed values match the expected result. Some easy prompts produce all-correct groups and some difficult prompts all-wrong groups, so the developer inspects how little comparative signal they supply. Training uses a more informative mixture. Evaluation on independently generated exercises checks answer correctness and whether the policy has learned parser-specific tricks rather than transferable arithmetic.
Limits and common mistakes
Relative normalization can interact with prompt difficulty and reward variance, while group sampling adds substantial generation cost. Identical rewards within a group provide weak or absent differentiation. A flawed verifier or biased reward model still produces flawed optimization. Length effects and normalization choices can influence which responses receive pressure. GRPO is neither a guarantee of reasoning nor a synonym for reinforcement learning from verifiable rewards. Evaluate the reward source, policy objective and task behavior as separate components.
Prerequisites
Related skills
- → is subcategory of: Reinforcement Learning
Sources and further reading
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Original group-relative policy optimization formulation.
- Hugging Face TRL: GRPO Trainer
Sampling groups, reward functions, normalization choices and training diagnostics.
Last updated: 2026-10-10