Atlas · skill

GRPO

Group Relative Policy Optimization updates a generating policy using rewards compared within groups of sampled responses. The skill is constructing informative prompt groups, implementing reliable rewards and controlling policy changes. GRPO can support verifiable or learned rewards; its optimizer is distinct from the source of the reward signal.

conceptAlignment

What it is

GRPO samples multiple completions for a prompt and uses their rewards to estimate relative advantages, commonly by centering and scaling scores within the group. It avoids a separately trained value model in the original formulation, using the group as its baseline instead. A clipped policy objective limits updates relative to the sampling policy, and formulations may include a penalty for divergence from a reference. Group size, generation settings and reward variability affect the learning signal. Modern implementations offer normalization and loss variants, so the exact objective needs documentation. GRPO does not define whether correctness, human preference or another property supplies the reward.

What the work involves

Test reward functions independently and select prompts where sampled completions exhibit meaningful quality differences. Choose group size and generation limits within the available inference and training budget. Inspect zero-variance groups, reward distributions, completion lengths, clipping and divergence. Keep problem families and solution sources separate from evaluation. Record the implementation's loss and normalization settings, then assess independently verified task accuracy rather than average reward alone. The useful result is a policy improvement with a reproducible sampling and optimization recipe and evidence that the relative signal corresponds to the intended behavior.

Illustrative example

An illustrative arithmetic policy generates several solutions for each exercise. A verifier rewards final answers only when parsed values match the expected result. Some easy prompts produce all-correct groups and some difficult prompts all-wrong groups, so the developer inspects how little comparative signal they supply. Training uses a more informative mixture. Evaluation on independently generated exercises checks answer correctness and whether the policy has learned parser-specific tricks rather than transferable arithmetic.

Limits and common mistakes

Relative normalization can interact with prompt difficulty and reward variance, while group sampling adds substantial generation cost. Identical rewards within a group provide weak or absent differentiation. A flawed verifier or biased reward model still produces flawed optimization. Length effects and normalization choices can influence which responses receive pressure. GRPO is neither a guarantee of reasoning nor a synonym for reinforcement learning from verifiable rewards. Evaluate the reward source, policy objective and task behavior as separate components.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10