---
type: paper
slug: additive-process-rewards
title: Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning
  — and What It Takes to Measure It
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Dmitri Kalaev
- Alexey Shvets
date: '2026-07-21'
status: AAAI 2027 submission
line: Verifiable reasoning
pages: 8
html: https://telegrapher.ai/research/additive-process-rewards/
pdf: https://telegrapher.ai/papers/additive-process-rewards/additive-process-rewards.pdf
reader: https://telegrapher.ai/research/additive-process-rewards/read/
json: https://telegrapher.ai/api/papers/additive-process-rewards.json
---

# Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It

## Paper gist

- **Claim:** In GRPO, an additive process reward's weight cancels if outcomes agree and is swamped if not
- **TL;DR:** Under GRPO's per-group standardization, a process term added to the outcome reward cannot be tuned: its weight cancels where a group's outcomes agree and is swamped where they differ. Endpoint gains need replication and a shuffle control.
- **Method:** Twelve GRPO runs of Qwen3.5-4B with QLoRA adapters (process weight γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each, group size G=4, 1,500 steps) were rewarded with a binary outcome plus a linter-based process score, evaluated greedily on the same 500 MATH-500 problems, and followed by a pre-registered fourteen-run probe of two fixes whose endpoint candidate went through a frozen fresh-seed and within-group-shuffle confirmation.
- **Key result:** Process-term displacement of standardized advantages in mixed-outcome groups: 0.3–6%; Training groups with zero advantage at γ = 0: 47.8%; Linter detection rate on injected errors: 99.5%
- **Why it matters:** Anyone adding a verifier score to a GRPO reward as an additive term gets a short design rule.
- **Limits:** Every run uses one model, Qwen3.5-4B with QLoRA adapters, on mathematical reasoning written in one compact trace language.
- **Status:** AAAI 2027 submission, July 2026
- **Read:** reader /research/additive-process-rewards/read/, PDF /papers/additive-process-rewards/additive-process-rewards.pdf

## Abstract

Verifiable process rewards—dense signals from checking intermediate reasoning steps—are widely expected to improve reinforcement learning of language-model reasoning. We show that in the standard group-normalized policyoptimization pipeline (GRPO), trajectory-additive process rewards are suppressed, in two regimes. In groups whose outcomes agree, the standardizing coefficient cancels exactly— an analytic consequence of the group-standard-deviation algebra—making the process term’s magnitude un-tunable, so coefficient sweeps are binary by construction. In groups whose outcomes disagree, the process term is compressed to a few percent of the advantage—empirically, a median of 0.3– 6% for process scores whose within-group spread is small relative to the outcome spread, as in our setting. Measuring this requires a verifier fast and deterministic enough to instrument every reward call: we use a compact machine-checkable trace language whose rule-based linter catches 99.5% of injected rule-family reasoning errors (vs. 66.3% for self-verification and 87.2% for a frontier LLM judge) at 9 ms per trace with no model in the loop—an operating point complementary to proof assistants, which verify formalized proofs rather than native traces. In a pre-registered 12-run sweep (4B model, mathematical reasoning), the mechanism’s predictions hold and endpoint effects are null within a ∼3pp detection floor. We derive and test two pre-registered fixes and an external-result arbitration; a separately pre-registered endpoint candidate that cleared multiple-comparison correction did not survive fresh-seed replication and a magnitude-matched semantic control.

## What we did and found

Take a group of four completions for one math problem, all of them wrong. GRPO subtracts the group's mean reward and divides by its standard deviation, so an outcome reward that is zero for every completion drops out of both, and the advantages come from the process score alone, divided by its own spread. Raise the process weight tenfold and numerator and denominator rise together. That case, and its opposite, where outcomes split and their spread sets the denominator, are the two regimes measured here for a trajectory-additive reward R = r_ans + γs. The process score s came from a rule-based linter for a compact trace language of givens, goals, tagged equations and explicit checks. It verifies a trace in 9 ms at the median, with a computer-algebra call and no model in the loop, which is fast enough to log group statistics at every reward call. Twelve Qwen3.5-4B policies were trained at γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each, and evaluated on the same 500 MATH-500 problems; a second pre-registered experiment of fourteen runs then tested what the account predicted.

In mixed-outcome groups the process term moved standardized advantages by a median of 0.3–6% across γ, while the outcome moved them by about one unit. In uniform-outcome groups γ cancels exactly, a corollary of the shared-denominator algebra that the paper does not claim as new. The two regimes interact. At γ = 0, 47.8% of training groups had zero advantage; any positive γ switched on those whose process scores differed, and the zero-advantage share fell to ≈32% at γ = 0.1, then stayed put through γ = 1.0. The sweep was close to a single on/off treatment, and its endpoints were null: the largest accuracy contrast, +1.47pp at γ = 1.0, had a 95% interval crossing zero. Dropping the standard-deviation divisor brought back a dose response at the advantage level, monotone but sub-linear, with six of six cell-seed observations inside the frozen bands. It also produced one endpoint candidate, +1.90pp verified-correct at Holm-adjusted p=0.006. On fresh seeds that gain was +0.00pp, and a within-group shuffled reward with the same per-group magnitudes reproduced and slightly exceeded the candidate's effect in every slice. A token-level alternative, at weight 0.5, padded traces with checkable lines and finished 7.5pp below baseline accuracy.

## Key numbers

| Measure | Value |
|---|---|
| Process-term displacement of standardized advantages in mixed-outcome groups | 0.3–6% |
| Training groups with zero advantage at γ = 0 | 47.8% |
| Linter detection rate on injected errors | 99.5% |
| Largest accuracy contrast in the 12-run sweep (γ = 1.0 vs γ = 0) | +1.47pp |
| Endpoint candidate's verified-correct gain on the selection set | +1.90pp |

## Why it matters

Anyone adding a verifier score to a GRPO reward as an additive term gets a short design rule. Under per-group standardization the coefficient is not a dial: it cancels wherever the process term is the only signal and is swamped wherever the outcome varies, so a sweep mostly tests whether the weight is above zero. A flat sweep is a null with a cause, not a verdict on process signals. For the signal to count at all, the standardization has to change (dropping the divisor restored dose response at the advantage level here) or the credit has to enter off the scalar advantage. That second route is not free. An uncentered per-token offset, paid whatever the outcome, became an annuity on checkable tokens and degraded the answers it was meant to support; centering and outcome-gating are the load-bearing choices.

The larger point is about evidence. The endpoint candidate was pre-registered, corrected for multiple comparisons, significant at p=0.006 and consistent with an independently confirmed mechanism. It still failed. The paper argues the exposure is general: any process reward adds reward mass that correlates with the metric it targets, and under group-relative advantages that mass shifts advantages whether or not it tracks anything semantic; held-out sets, extra training seeds and multiple-comparison correction all leave such a gain looking semantic. A within-group permutation of the process reward keeps its magnitude distribution and breaks its link to trace content, which separates the two. The paper proposes it as a standard, near-zero-cost control for endpoint claims about verifiable process rewards.

The trace language and its linter come from Telegraph Reasoning, which introduces a grammar that lets a rule-based linter and a symbolic algebra system check every step of a trace; here they are the measuring instrument rather than the result. The same language restricts what a model may write, and so bounds accuracy apart from any reward. Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning takes up that cost, scoring checker acceptance and answer accuracy on the same outputs against free-form reasoning, and finds that the two come apart.

## What this does not show

Every run uses one model, Qwen3.5-4B with QLoRA adapters, on mathematical reasoning written in one compact trace language. The cancellation in uniform-outcome groups follows from the algebra of any group-standardized objective, but the few-percent displacement in mixed groups is an empirical magnitude: it depends on process scores varying much less within a group than outcomes do, and it was measured at one scale, one group size (G=4) and one process-score scale, so larger models or other verifiers may sit elsewhere. The endpoint statistics are powered for effects of roughly three percentage points. The sweep is a null within that floor, and the γ = 1.0 accuracy lean, positive in three of three seeds, is unproven rather than disproven; the pooled problem-seed test also treats each checkpoint as fixed, so its intervals are narrower than a seed-random-effects analysis would give. The linter's detection comparison is not false-positive-matched and draws its errors from the linter's own rule vocabulary. The trace language bounds accuracy independently of the reward, and the paper does not disentangle the two. Reading the sub-linear dose response as self-damping is post hoc. The arbitration of an external positive result with a deterministic step verifier is inconclusive: the largest-share configuration did not reproduce that benefit, but group size, reward aggregation and domain differ, and the external configuration was not run at this scale. And the shuffle control rejects a semantic reading of the candidate's gain without saying whether seed selection, optimization noise or a magnitude effect produced it.

## Blog post

[In GRPO, the weight on an additive process reward acts like a switch, not a dial](https://telegrapher.ai/blog/additive-process-rewards.md)

## Cite

```bibtex
@misc{bei2026why,
  title         = {Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
  year          = {2026},
  note          = {AAAI 2027 submission},
  url           = {https://telegrapher.ai/research/additive-process-rewards/}
}
```
