telegrapher

In GRPO, an additive process reward's weight cancels if outcomes agree and is swamped if not

Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Dmitri Kalaev, Alexey Shvets

AAAI 2027 submission, July 2026

Under GRPO's per-group standardization, a process term added to the outcome reward cannot be tuned: its weight cancels where a group's outcomes agree and is swamped where they differ. Endpoint gains need replication and a shuffle control.

What we did and found

Take a group of four completions for one math problem, all of them wrong. GRPO subtracts the group's mean reward and divides by its standard deviation, so an outcome reward that is zero for every completion drops out of both, and the advantages come from the process score alone, divided by its own spread. Raise the process weight tenfold and numerator and denominator rise together. That case, and its opposite, where outcomes split and their spread sets the denominator, are the two regimes measured here for a trajectory-additive reward R = r_ans + γs. The process score s came from a rule-based linter for a compact trace language of givens, goals, tagged equations and explicit checks. It verifies a trace in 9 ms at the median, with a computer-algebra call and no model in the loop, which is fast enough to log group statistics at every reward call. Twelve Qwen3.5-4B policies were trained at γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each, and evaluated on the same 500 MATH-500 problems; a second pre-registered experiment of fourteen runs then tested what the account predicted.

In mixed-outcome groups the process term moved standardized advantages by a median of 0.3–6% across γ, while the outcome moved them by about one unit. In uniform-outcome groups γ cancels exactly, a corollary of the shared-denominator algebra that the paper does not claim as new. The two regimes interact. At γ = 0, 47.8% of training groups had zero advantage; any positive γ switched on those whose process scores differed, and the zero-advantage share fell to ≈32% at γ = 0.1, then stayed put through γ = 1.0. The sweep was close to a single on/off treatment, and its endpoints were null: the largest accuracy contrast, +1.47pp at γ = 1.0, had a 95% interval crossing zero. Dropping the standard-deviation divisor brought back a dose response at the advantage level, monotone but sub-linear, with six of six cell-seed observations inside the frozen bands. It also produced one endpoint candidate, +1.90pp verified-correct at Holm-adjusted p=0.006. On fresh seeds that gain was +0.00pp, and a within-group shuffled reward with the same per-group magnitudes reproduced and slightly exceeded the candidate's effect in every slice. A token-level alternative, at weight 0.5, padded traces with checkable lines and finished 7.5pp below baseline accuracy.

Key numbers

Process-term displacement of standardized advantages in mixed-outcome groupsmedian across γ, groups with σ_R ≥ 0.05; outcome-driven displacements are on the order of one unit0.3–6%
Training groups with zero advantage at γ = 018.0% of all groups carry process scores that could separate them; any γ > 0 cuts the zero-advantage share to ≈32%, unchanged up to γ = 1.047.8%
Linter detection rate on injected errorsat a 10% false-positive rate, 9 ms per trace; self-verification 66.3%, frontier LLM judge 87.2%; rates not false-positive-matched and errors drawn from the linter's own rule vocabulary99.5%
Largest accuracy contrast in the 12-run sweep (γ = 1.0 vs γ = 0)paired 95% CI [−0.27, +3.20], McNemar p = 0.115 over 1,500 paired problem-seed observations (500 MATH-500 problems, three seeds); design powered for roughly 3pp+1.47pp
Endpoint candidate's verified-correct gain on the selection setHolm-adjusted p=0.006 at the 256-token cap; +0.00pp on fresh seeds, and a magnitude-preserving shuffled reward reproduced and slightly exceeded the candidate's effect in every slice+1.90pp

What this does not show

Every run uses one model, Qwen3.5-4B with QLoRA adapters, on mathematical reasoning written in one compact trace language. The cancellation in uniform-outcome groups follows from the algebra of any group-standardized objective, but the few-percent displacement in mixed groups is an empirical magnitude: it depends on process scores varying much less within a group than outcomes do, and it was measured at one scale, one group size (G=4) and one process-score scale, so larger models or other verifiers may sit elsewhere. The endpoint statistics are powered for effects of roughly three percentage points. The sweep is a null within that floor, and the γ = 1.0 accuracy lean, positive in three of three seeds, is unproven rather than disproven; the pooled problem-seed test also treats each checkpoint as fixed, so its intervals are narrower than a seed-random-effects analysis would give. The linter's detection comparison is not false-positive-matched and draws its errors from the linter's own rule vocabulary. The trace language bounds accuracy independently of the reward, and the paper does not disentangle the two. Reading the sub-linear dose response as self-damping is post hoc. The arbitration of an external positive result with a deterministic step verifier is inconclusive: the largest-share configuration did not reproduce that benefit, but group size, reward aggregation and domain differ, and the external configuration was not run at this scale. And the shuffle control rejects a semantic reading of the candidate's gain without saying whether seed selection, optimization noise or a magnitude effect produced it.

Every number above was checked against the paper text.
The authors' abstract

Verifiable process rewards—dense signals from checking intermediate reasoning steps—are widely expected to improve reinforcement learning of language-model reasoning. We show that in the standard group-normalized policyoptimization pipeline (GRPO), trajectory-additive process rewards are suppressed, in two regimes. In groups whose outcomes agree, the standardizing coefficient cancels exactly— an analytic consequence of the group-standard-deviation algebra—making the process term’s magnitude un-tunable, so coefficient sweeps are binary by construction. In groups whose outcomes disagree, the process term is compressed to a few percent of the advantage—empirically, a median of 0.3– 6% for process scores whose within-group spread is small relative to the outcome spread, as in our setting. Measuring this requires a verifier fast and deterministic enough to instrument every reward call: we use a compact machine-checkable trace language whose rule-based linter catches 99.5% of injected rule-family reasoning errors (vs. 66.3% for self-verification and 87.2% for a frontier LLM judge) at 9 ms per trace with no model in the loop—an operating point complementary to proof assistants, which verify formalized proofs rather than native traces. In a pre-registered 12-run sweep (4B model, mathematical reasoning), the mechanism’s predictions hold and endpoint effects are null within a ∼3pp detection floor. We derive and test two pre-registered fixes and an external-result arbitration; a separately pre-registered endpoint candidate that cleared multiple-comparison correction did not survive fresh-seed replication and a magnitude-matched semantic control.

On the blogIn GRPO, the weight on an additive process reward acts like a switch, not a dialGroup standardization cancels an additive process reward's weight when outcomes agree and swamps the term when they differ. We measured both, then tested fixes.

More in Verifiable reasoning

Cite

@misc{bei2026why,
  title         = {Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
  year          = {2026},
  note          = {AAAI 2027 submission},
  url           = {https://telegrapher.ai/research/additive-process-rewards/}
}

Builds on

  1. Shao et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.
  2. Liu et al. (2025). Understanding R1-Zero-Like Training: A Critical Perspective.
  3. Bay and Yearick (2026). GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity.
  4. Pronesti et al. (2026). Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning.
  5. Wang et al. (2026). Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners.