In GRPO, the weight on an additive process reward acts like a switch, not a dial
Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure ItThe paper: summary, reader and PDF
You are training a reasoning model with GRPO on math problems with checkable answers, and you have a checker that can verify individual lines of a derivation. Rewarding derivations that check out, on top of correct answers, looks like an obvious extra signal. So you add a process term to the reward, pick a weight, and sweep it.
We ran that sweep, pre-registered: twelve runs, four weights, three seeds each. Accuracy did not move beyond noise at any weight. The tempting conclusion is that process rewards do not help. The data support a narrower one: under GRPO’s group standardization, the weight we swept had little say over how much process signal reached the policy. The paper works out why, and what it took to measure.
Four completions, two ways to lose the dial
GRPO samples a group of completions per prompt (four, in our runs), subtracts the group’s mean reward from each and divides by the group’s standard deviation, giving the advantage each completion is trained on. Our reward was the outcome, one if the answer is right and zero if not, plus a weight γ times a process score from the checker.
Suppose all four completions are wrong. The outcome is identical across the group, so it drops out of both the mean and the standard deviation. What remains is the process score divided by its own spread, and γ multiplies the top and bottom of that fraction alike: set it to 0.1 or to 1.0 and the advantages come out the same. The process term is now the whole advantage, at full strength. What it has lost is the dial.
Now suppose two are right and two are wrong. The right-versus-wrong gap dominates the standard deviation, and the process score, which varies far less within a group, shrinks to a sliver. Across our sweep it shifted advantages by a median of 0.3–6%; the outcome moved them by about one unit.
The paper calls these the two regimes of washout, and they differ in kind. The all-wrong case is pure algebra — a corollary, as the paper presents it, of the shared-denominator identity that recent GRPO analyses already use, not a new result. The mixed case is a measured magnitude that depends on our process scores varying little within groups.
Together they produce something awkward. At γ = 0, 47.8% of training groups had identically zero advantage and gave no gradient, and more than a third of those carried process scores that differed. Any positive weight switches those groups on, and the switch saturates: by γ = 0.1 the zero-advantage share had fallen to roughly a third of groups, and it did not move as γ rose tenfold to 1.0. A sweep over γ mostly tests whether γ is above zero.
A checker fast enough for the reward loop
Seeing this meant scoring every completion at every reward call. That rules out learned process-reward models, which are slow models themselves, and proof assistants, which certify far more but need formalized statements and take seconds to minutes.
We used the trace language from Telegraph Reasoning. The model writes givens, goals, tagged equations and explicit checks in a fixed line grammar, and a rule-based linter with a computer-algebra call marks each equation and check line pass, fail or uncheckable. No model runs inside the checker. A median trace takes about nine milliseconds. On a benchmark of injected errors it caught 99.5%, against roughly two thirds for self-verification and under nine in ten for a frontier LLM judge. That comparison flatters the linter: false-positive rates are unmatched, and the errors come from the vocabulary its rules encode. The firmer evidence is that on organic output its verdicts varied across runs and moved over training. It checks line-local consistency, not theoremhood, which is enough for this measurement.
The sweep came back flat
Twelve Qwen3.5-4B policies with QLoRA adapters, 1,500 optimization steps each, were evaluated greedily on the same 500 MATH-500 problems. Means over three seeds:
| Process weight γ | Accuracy (%) | Change vs γ = 0, pp [95% CI] | Linter-valid traces (%) |
|---|---|---|---|
| 0 (outcome only) | 37.07 | — | 23.53 |
| 0.1 | 37.67 | +0.60 [−1.20, +2.40] | 25.73 |
| 0.5 | 36.93 | −0.13 [−1.87, +1.67] | 25.47 |
| 1.0 | 38.53 | +1.47 [−0.27, +3.20] | 23.67 |
No interval excludes zero. The design was powered for effects of roughly three percentage points and cannot rule out smaller ones, so this is a null within that floor; the lean at γ = 1.0, positive in all three seeds, is unproven rather than disproven. The low-γ bump in linter validity does not survive correction across the primary test family.
Two routes around the standardization
A second pre-registered experiment, fourteen runs with numeric bands frozen in advance, fixed γ at 1 and varied β, the weight on the line-credit part of the process score. The prediction: switch off the standard-deviation division (scale_rewards=none, keeping the mean subtraction) and the process term’s share of the advantage grows with β, within a factor of two of linear. All six cell-seed observations landed inside the bands. The rise was monotone but sub-linear, the share growing a little under sixfold while β grew tenfold. After the fact, we read that as self-damping: a large process term inflating the group spread it is normalized against. In uniformly wrong groups, advantages now scaled with β instead of saturating near ±1. The dial was back.
Token-level credit, the other route, broke legibly. Our version gave the tokens inside each line the linter judged a fixed signed offset from its verdict on that line, uncentered and blind to the outcome. At weight 0.1 this was a mild, harmless drift. At 0.5 the policy wrote about a fifth fewer characters but roughly three fifths more equation lines and three quarters more check lines. Held-out accuracy slid over training in both seeds, and final accuracy ended well below baseline. The reward had become an annuity on checkable-line tokens. That rules out this design at meaningful weight, not token-level credit as such: centering and outcome-gating are what such credit has to get right.
A significant result that did not survive its controls
Restoring the dial produced one endpoint candidate. With the divisor removed and β = 0.5, the verified-correct rate (answer right and trace linter-valid) rose 1.90 points on the selection set at the tightest token cap, Holm-adjusted p = 0.006. Pre-registered, corrected, significant, mechanism-consistent: by most standards, a finding.
Under a confirmation protocol fixed in advance, it failed twice. On fresh seeds the gain was +0.00pp. And a shuffled control, which permutes the process reward within each group so its magnitudes stay while its link to trace content goes, reproduced and slightly exceeded the candidate’s effect in every slice.
We argue the problem is general rather than a quirk of our setup. Any process reward adds reward mass that correlates with the metric it targets, and under group-relative advantages that mass shifts advantages whether or not it tracks anything semantic. The endpoint then moves the way a real effect would. Held-out sets, multiple training seeds and multiple-comparison correction leave that confound standing; a within-group shuffle does not.
What changes for anyone adding a process reward
Under per-group standardization, the coefficient on an additive process term will not tune it. For the signal to count at all, the standardization has to change, or the credit has to enter off the scalar advantage, where centering and outcome-gating carry the load. Neither route produced an endpoint gain here that survived replication and the shuffle.
Before trusting an endpoint gain, rerun it on fresh seeds and with the process reward shuffled within groups. The shuffle costs next to nothing, and the paper proposes it as a standard control.
What is still open
Everything here comes from one 4B model, on mathematics, in one trace language, at one group size. The cancellation holds for any group-standardized objective; the few-percent displacement is a measured magnitude that larger models, other group sizes or another verifier could shift.
The trace language has a cost of its own. It restricts what the model can write, which bounds accuracy whatever the reward, and this paper does not separate the two effects. Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning takes up that cost: scoring checker acceptance and answer accuracy on the same outputs against free-form reasoning, it finds that the trace interface raises acceptance and lowers accuracy.
The shuffle, for its part, rejects a semantic reading of the candidate’s gain without saying what did produce it: seed selection, optimization noise and a magnitude effect all remain possible.
The mechanism made one endpoint prediction, and it failed. A deterministic step verifier used as an additive reward has been reported to improve a structured-reasoning task (Pronesti, Belz and Hou, 2026), and our mechanism’s simplest reading said our largest-share configuration should show the endpoint effect smaller ones lacked. It did not. Their setting differs in group size, per-step versus per-trajectory reward, and domain, and we have not run it at our scale, so we count our prediction as failed, not their result as explained.
Restoring the dial was the easy part. We have not yet found a setting where turning it moves accuracy on fresh seeds and stops moving it once the reward is shuffled.