# Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It

Full text, page by page. Paper page: https://telegrapher.ai/research/additive-process-rewards.md

## Page 1

Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement
Learning — and What It Takes to Measure It

Anonymous submission

Abstract

Verifiable process rewards—dense signals from checking
intermediate reasoning steps—are widely expected to improve reinforcement learning of language-model reasoning.
We show that in the standard group-normalized policyoptimization pipeline (GRPO), trajectory-additive process rewards are suppressed, in two regimes. In groups whose outcomes agree, the standardizing coefficient cancels exactly—
an analytic consequence of the group-standard-deviation
algebra—making the process term’s magnitude un-tunable,
so coefficient sweeps are binary by construction. In groups
whose outcomes disagree, the process term is compressed to
a few percent of the advantage—empirically, a median of 0.3–
6% for process scores whose within-group spread is small relative to the outcome spread, as in our setting. Measuring this
requires a verifier fast and deterministic enough to instrument
every reward call: we use a compact machine-checkable trace
language whose rule-based linter catches 99.5% of injected
rule-family reasoning errors (vs. 66.3% for self-verification
and 87.2% for a frontier LLM judge) at 9 ms per trace with
no model in the loop—an operating point complementary to
proof assistants, which verify formalized proofs rather than
native traces. In a pre-registered 12-run sweep (4B model,
mathematical reasoning), the mechanism’s predictions hold
and endpoint effects are null within a ∼3pp detection floor. We
derive and test two pre-registered fixes and an external-result
arbitration; a separately pre-registered endpoint candidate that
cleared multiple-comparison correction did not survive freshseed replication and a magnitude-matched semantic control.

1

Introduction

Reinforcement learning with verifiable rewards has become
the dominant recipe for improving the reasoning of language
models, and a natural extension is to reward not just the final
answer but the intermediate steps: if a checker can verify
each line of a derivation, why not shape the policy toward
derivations that check out? This intuition motivates a large
and active body of process-reward and step-verification work.
We report a mechanism that undercuts the simplest version
of it, an instrument built to measure that mechanism, and a
program of pre-registered experiments, one of which shows
concretely what “measuring” an endpoint effect requires.

This is an anonymized submission for review purposes only. Distribution, citation, or public sharing of this manuscript is strictly
prohibited. Copyright and publication details will appear in the
final version if accepted.

The mechanism is a property of group-relative policy optimization (GRPO), the standard critic-free algorithm, and
specifically of its per-group reward standardization. When
the reward is a trajectory-additive composite of a binary outcome term and a bounded process term, the process term’s
influence on the policy gradient splits into two regimes with
opposite pathologies. In sampling groups whose completions disagree on the outcome, the outcome variance dominates the group’s standard deviation and the process term
is compressed to a few percent of the standardized advantage (a median of 0.3 to 6% across our sweep). In groups
whose completions agree on the outcome, the standardizing
denominator becomes the process term’s own spread, the coefficient cancels exactly, and the term’s magnitude becomes
untunable, so a coefficient sweep is a binary experiment in
disguise. Neither regime is a bug in an implementation; both
follow from the standardization algebra, and together they
retrodict a clean null we measured earlier and predict what
happens when the standardization is changed.
Measuring any of this requires a verifier fast and deterministic enough to score every completion at every reward call —
tens of thousands of line-level checks per run — which rules
out both the learned process-reward models that dominate the
literature (they are themselves models, and slow) and the formal proof assistants that anchor its rigorous end (they verify
formalized statements behind an autoformalization barrier,
at seconds-to-minutes latency). We use a compact, machinecheckable trace language whose rule-based linter verifies a
native reasoning trace in about nine milliseconds with no
model in the loop, detecting 99.5% of injected reasoning
errors against 66.3% for self-verification and 87.2% for a
frontier LLM judge. The linter is not the contribution; it is
the instrument that makes the mechanism visible, and its operating point — deterministic verification of native traces,
inside the reward loop — is complementary to, not competitive with, the proof-assistant end of the spectrum.
The paper makes three contributions. The two-regime
washout mechanism comes with quantitative predictions that
we confirm out of sample: disabling the standardizing divisor restores dose-responsive process influence across a coefficient ladder (six of six seed observations inside pre-registered
bands), while an alternative token-level injection path breaks
accuracy outright in a fully attributable way, which marks
centering and outcome-gating as the load-bearing design

## Page 2

choices. The instrument and its detection study are the second contribution, situating deterministic native-trace verification against proof assistants and learned reward models.
The methodological core is a pre-registered endpoint experiment in which the single surviving candidate — significant after multiple-comparison correction (p=0.006) and
consistent with the mechanism — failed to replicate on
fresh seeds and failed a magnitude-preserving, semanticsdestroying control whose verdict we had committed to in
advance. At the scale we study (a 4B model on mathematical
reasoning), no process-reward configuration shows an endpoint benefit that survives replication and a semantic control,
in either injection path. The episode answers the question in
our title: measuring a verifiable-reasoning reward’s effect
takes more than a clean gate; it takes replication and a control that separates what the verifier measures from the reward
mass it adds.

2

Background and Related Work

Reasoning with and about intermediate steps. Chainof-thought prompting established that eliciting intermediate
steps improves mathematical reasoning (Wei et al. 2022), on
benchmarks that remain the field’s yardsticks (Cobbe et al.
2021; Hendrycks et al. 2021). Reinforcement learning with
verifiable rewards then made the final answer the training
signal: group-relative policy optimization (GRPO) replaced
the learned critic with within-group reward standardization
(Shao et al. 2024), and scaled outcome-only training to strong
reasoning models (DeepSeek-AI 2025). Our setting is exactly
this pipeline, with one addition, a process term in the reward,
and our subject is what the pipeline itself does to that addition.

What group normalization does. A rapidly consolidating line analyzes GRPO’s normalization for pure outcome
rewards: removing the standard-deviation division to correct a difficulty bias (Liu et al. 2025), decoupled clipping
and dynamic sampling (Yu et al. 2025), curvature accounts
of why dividing by the group deviation helps (Ge et al.
2026), pathologies of the z-score under low-dispersion rewards (Salmani-Zarchi et al. 2026), corrections to the groupmean baseline (Garg et al. 2025), methods for exploiting
the zero-variance groups that vanilla GRPO discards (Le
et al. 2025), and a unification showing the major variants act
on a single quantity, the group standard deviation (Bay and
Yearick 2026). All of these treat rewards as given, and binary. To our knowledge none analyzes the composite case (an
additive process term riding on the outcome reward), where,
as Section 4 shows, the same denominator that stabilizes
outcome learning silently reallocates and caps the shaping
signal.

Process rewards and fine-grained credit. Process supervision at the step or token level is an active fix-space: aligning
process with outcome rewards in policy optimization (Ding
et al. 2026), generative credit assignment (Xie et al. 2025),
temporal-difference-style credit in GRPO (Parthasarathi et al.
2025), outcome-grounded advantage reshaping (Li et al.
2026), and token-level reward models (Chen et al. 2025).

These works propose injection schemes. Contemporaneously, PASS (Wang et al. 2026a) observes that a composite
reward’s process channel can be contaminated under grouprelative optimization; our contribution is distinct and complementary along four axes — a two-regime measurement
of the advantage displacement, the analytic cancellation of
the additive coefficient in zero-outcome-variance groups, a
deterministic verifier fast enough to instrument every reward call, and a semantic-shuffle-plus-fresh-seed discipline
for endpoint claims — and Section 6 uses the measurement
to make part of this design space testable at the advantage
level (our confirmed predictions there are advantage-level,
not endpoint-level). Closest to our arbitration experiment, a
deterministic rule-based step verifier has been reported to
improve a domain-specific reasoning task as an additive reward under this same family of pipelines (Pronesti, Belz, and
Hou 2026); our mechanism’s simplest reading places that result in the regime where the aggregate process term is large
relative to the outcome term, and we probe the reconciliation
directly. Reward-hacking analyses motivate our monitoring
choices (Wang et al. 2026b).

Verifying reasoning. Approaches to checking model reasoning range from formal proof assistants, which certify
formalized statements at seconds-to-minutes latency behind
an autoformalization barrier, through learned process-reward
models and LLM judges, to deterministic symbolic checkers applied at evaluation time (He et al. 2026; Su 2025;
Zhou et al. 2026). Section 3 argues these are operating points
on a latency–coverage–trust frontier, and that a deterministic, model-free checker over native traces is the operating
point the reward loop of reinforcement learning can afford.
Efficient- reasoning work (length penalties, budget control,
the accuracy-per-token trade-off; Sui et al. 2025; Dumitru
et al. 2025; Yi and Wang 2025; Wen et al. 2025; Zhang et al.
2026; Gao et al. 2025) is adjacent but orthogonal here: we
use token budgets only as evaluation protocol, and make no
efficiency claim in this paper.

3

The Instrument: A Deterministic Trace
Verifier

Measuring Section 4’s quantities requires scoring the process quality of every completion at every reward call—and
doing so deterministically, because a verifier whose own
judgments are noisy or model-dependent cannot serve as
the measurement standard for an analysis of reward noise.
We use Telegraph Reasoning (TE), a compact trace language
in which the model states givens, goals, tagged equations,
and explicit checks in a fixed line grammar, terminating in a
machine-parsable answer line. A rule-based linter enforces
the grammar and, for each equation or check line, attempts
deterministic verification by symbolic substitution: every line
receives a verdict of pass, fail, or uncheckable, and the trace
as a whole a structural validity bit. Twelve rules cover structure, scoping, and numeric consistency; the implementation
is pure rules plus a computer-algebra call, with no language
model anywhere in the loop, and verifies a trace in 9 ms at
the median. The full grammar and rule specification appear
in the supplementary material.

## Page 3

Injected errors

Organic traces

Verifier

TPR

FPR

Loc.±3

valid-rate span

TE linter (ours)
Self-verify (NP)
LLM judge (frontier)

99.5
66.3
87.2

10.0
—
—

90.3
—
—

21.8–27.2
n/a
n/a

Table 1: Discriminative power of the instrument (Section 3).
Injected-error columns from the corruption benchmark (five
corruption types; localization = fraction of detections within
±3 lines); the benchmark favors the linter by construction
(shared rule vocabulary), so the final column reports discrimination on organic RL outputs: full-trace validity span
across the twelve γ-sweep cells.

Why deterministic: discriminative power. On a corruption benchmark of reasoning traces with injected errors spanning five corruption types, the linter detects 99.5% of errors
at a 10% false-positive rate; a self-verification protocol in
the Natural Program style detects 66.3%, and a frontier LLM
judge prompted for the same task 87.2%. Two limitations
qualify this comparison. The detection rates compare truepositive rates that are not false-positive-matched across the
three verifiers, so the headline gap overstates a like-for-like
advantage; and the injected errors are drawn from the same
failure vocabulary the linter’s rules encode, so the benchmark is favorable to it. Two observations temper the concern. First, the same study measures localization—90.3% of
detections within three lines of the injected fault—which a
vocabulary-matched but shallow checker would not achieve.
Second, and more directly, the linter discriminates on organic
model output, where no injection occurred: across our twelve
reinforcement-learning runs its full-trace validity rate spans
21.8–27.2% by cell on the evaluation set, moves measurably
over training on held-out problems (by −2.0 to +8.0pp per
run), and its per-line verdicts generate the group statistics of
Section 4—structure that could not arise from a checker that
merely echoed its own injections.

The operating point: verification inside the reward loop.
Formal proof assistants occupy one end of the verification
spectrum: they certify complete correctness, but of formalized statements, and reaching them from natural mathematical text requires autoformalization, which remains unreliable
precisely on the informal reasoning reinforcement learning
must score, with verification latencies of seconds to minutes
per attempt (DeepSeek-AI 2025; Pronesti, Belz, and Hou
2026). Learned process-reward models occupy the other end:
they score native text directly, but are themselves models—
opaque, biased toward surface features, and attackable by
the very policy they train (Wang et al. 2026b). The linter
sits at a third point: verification of native traces, deterministic, and cheap enough to be invisible in the training budget.
A single training run here makes roughly 9×10 3 full-trace
verifications (6,000 in the reward path, 3,000 on held-out
checkpoints) comprising some 4×10 4 line-level checks; at
9 ms per trace this is well under 0.1% of a ∼30-second optimization step, where a seconds-scale verifier would consume
10–40% of it and an autoformalization round-trip is infeasi-

ble in-loop. The costs are complementary rather than competitive: the linter certifies much less than a proof assistant,
checking line-local consistency rather than theoremhood, but
it is a practical operating point that can sit inside the reward
loop and stamp every completion the policy produces, which
is what the measurement in this paper requires. The trace language restricts what the model may write, and we measure
rather than assume the cost of that restriction: expressibility coverage of the training distribution is reported in the
supplementary material.

4

The Mechanism: Two Regimes of Washout

Group-relative policy optimization draws G completions
per prompt, scores each with a scalar reward, and converts
rewards to advantages by standardizing within the group:
A i = (R i − R̄)/(σ R + ϵ), with σ R the within-group sample
standard deviation and ϵ = 10 −4 (Shao et al. 2024). When the
reward is a trajectory-additive composite R i = r ans,i + γ s i ,
where r ans ∈ {0, 1} is the verifiable outcome and s i is a
bounded process score, the fate of the process term is decided entirely by the composition of the group, and it splits
into two regimes with opposite pathologies.

Regime 1: outcome variance present. If the group contains both correct and incorrect completions, σ R is dominated by the outcome spread. In our training runs the withingroup standard deviation of r ans averages 0.19, while that
of γs reaches only 0.002 at γ=0.1 and 0.021 at γ=1.0 with
β=0.1 process scoring (Section 5). Restricted to groups with
a well-conditioned denominator (σ R ≥ 0.05), the process
term displaces post-normalization advantages by a median
of 0.3–6% across γ, against outcome-driven displacements
on the order of one unit (the full per-γ series is in the supplement). This is an empirical magnitude bound conditional
on σ(γs) ≪ σ(r ans ) — a fact about our process-score construction, not the standardization algebra — measured at one
scale, group size (G=4), and model class; it is not an impossibility theorem.

Regime 2: outcome variance absent. If every completion
in the group receives the same outcome score—all wrong,
or all right—the outcome contributes nothing to either the
numerator or the denominator, and the advantage reduces to
A i = γ(s i − s̄)/(γσ s + ϵ). For γσ s ≫ ϵ the coefficient
cancels:

Lemma 1. In a group with constant outcome reward and
non-degenerate process scores, group-standardized advantages are invariant to the process coefficient: A i → (s i −
s̄)/σ s as γσ s /ϵ → ∞, independent of γ.

The lemma is an immediate consequence of the shareddenominator algebra that recent analyses have used to unify
GRPO variants (Bay and Yearick 2026), and we present it
as such rather than as a novelty; its practical force in the
composite-reward setting appears to be unexamined. In these
groups the process term is not suppressed at all: it is the entire
standardized advantage, at full magnitude, but its strength is
untunable, since any γ > 0 produces the same advantages as
any other.

## Page 4

The rescue channel and its saturation. Regime 2 is not
a corner case. At γ = 0, 47.8% of training groups produce
identically zero advantage (no gradient from the prompt at
that step), and 18.0% of all groups — roughly 38% of the
zero-variance ones — carry non-degenerate process scores
that could differentiate their completions. Introducing any
positive coefficient activates every such group and then saturates: the zero-advantage fraction drops to ≈32% at γ = 0.1
and does not move further as γ grows to 1.0 (a matching
saturation appears in per-group ranking statistics, reported
in the supplement). The consequence for experiment design
is a scoped one: because the live rescue channel saturates
at any γ > 0 while the Regime-1 scaling channel remains
but is bounded at the measured displacement share, a coefficient sweep of a trajectory-additive process reward behaves
as an approximately binary treatment (1[γ > 0]) rather than
a dose. The two regimes are asymmetric in status: Regime 2
is analytic (Lemma 1, a corollary of the shared-denominator
algebra), while Regime 1 is a purely empirical magnitude observation. This picture retrodicts the shape of our endpoint
results (Section 5)—no dose-response was constructible—
and generates the out-of-sample predictions of Section 6:
removing the standard-deviation division should restore βlinear process influence in all groups (prediction D1); injecting credit at the token level bypasses the scalar bottleneck
entirely (D3); and additive composites in which the aggregate process term is large relative to the outcome term—the
regime of a recently reported positive result with a deterministic step verifier (Pronesti, Belz, and Hou 2026)—should
escape Regime 1’s bound (D4).
These within-group statistics are logged at every reward
call (Section 3 describes the verifier that makes this practical)
and are stationary across training. One measurement caveat
is load-bearing: the ϵ-regularized denominator inflates naive
means over displacement ratios by 20–30× through a small-
σ R tail, so we report medians and restricted-denominator
variants throughout.

5

A Pre-Registered Null, Explained

A note on notation before the experiments. Two coefficients
appear: γ scales the entire process bundle s in the trajectory
reward, while β scales only the line-credit term inside s.
The sweep of Section 5 varies γ at fixed β=0.1; the probe
of Section 6 instead fixes γ=1 and varies β, so “the process
weight” there refers to β. Where Section 4’s cancellation is at
issue, it is γ (the additive-bundle coefficient) that cancels, independent of β. Because β leaves the uncheckability penalty
fixed while γ scales the complete bundle, the two sweeps test
related but non-equivalent interventions.
We trained twelve policies—γ ∈ {0, 0.1, 0.5, 1.0}, three
seeds each—on a shared pool of mathematics problems disjoint from all evaluation sets, under a single hardware and
kernel configuration, with the γ = 0 cells re-run inside the
sweep so that the outcome-only baseline shares every execution detail with the treatment arms. Concretely: a 4B model
(Qwen3.5-4B) with QLoRA adapters, group size G=4, 1,500
optimization steps per run, greedy decoding, evaluated at
generation caps {256, 512} and full length. The process score
s = βr line − p uncheck bundles the linter’s per-line pass rate

∆Acc vs γ=0
paired [95% CI]

γ

Accuracy
mean ± sd

Lint-valid
mean ± sd

Verif.-
corr.

0.0
0.1
0.5
1.0

37.07 ± 0.90
37.67 ± 0.76
36.93 ± 0.70
38.53 ± 0.31

23.53 ± 1.45
25.73 ± 1.29
25.47 ± 1.10
23.67 ± 2.01

13.93
—
14.73 +0.60 [−1.20, +2.40]
14.33 −0.13 [−1.87, +1.67]
14.00 +1.47 [−0.27, +3.20]

Table 2: The γ-sweep endpoints (Section 5). Twelve runs,
uniform configuration; γ=0 is the outcome-only baseline
re-run inside the sweep. No paired contrast excludes zero
(largest: γ=1.0, McNemar p=0.115); the lint-valid low-γ
bump does not survive family-wise correction (Section 5).
Verified-correct = answer correct and trace linter-valid.

with an uncheckability penalty at fixed β = 0.1; γ = 0
reproduces the outcome-only reward bit-for-bit, a property
locked by unit test rather than assumed. All twelve cells were
evaluated greedily on the same 500 MATH-500 problems
in identical order, enabling paired-by-problem inference; the
gate criteria, sweep grid, and diagnostics were fixed before
any training run.
Accuracy shows no effect at any coefficient (Table 2).
The largest contrast, γ = 1.0 against γ = 0, is +1.47pp
with a paired 95% confidence interval of [−0.27, +3.20] and
McNemar p = 0.115 over 1,500 paired problem-seed observations. This pooled problem-seed test is conditional on the
realized policies: it treats each trained checkpoint as fixed
and does not estimate the additional uncertainty over training seeds, so its intervals are narrower than a seed-random-effects analysis would give (a limitation we return to for the
confirmation stage). The per-problem dose-response slope
is +1.11pp per unit γ with an interval crossing zero, and
a label-permutation test over the twelve cell means gives
p = 0.095. Linter-validity, the metric the process term directly rewards, shows a suggestive low-coefficient anomaly
(+2.20pp at γ=0.1, uncorrected p = 0.014; +1.93pp at
γ=0.5) that is absent at γ = 1.0, survives Holm correction only within its own three-test family, dies across the
nine-test primary family, and vanishes under the pooled any-
γ-positive contrast (+1.42pp, CI [−0.07, +2.93]). Verifiedcorrect rate (answer correct and trace linter-valid) is null
everywhere. Across the seventeen tests this analysis reports,
roughly 0.8 would be expected below p = 0.05 under a
global null; exactly two were, both in the linter-validity family just described. For the program as a whole, we evaluated
two pre-registered endpoint gates (this γ-sweep and the coefficient/injection probe of Section 6); the single candidate
that cleared either was carried to a dedicated confirmation
stage whose analysis was frozen before it ran (Section 6), so
the search over configurations that produced it is handled by
replication rather than by a family-wise correction over the
whole program.
Two qualifications are part of the claim rather than caveats
to it. First, the design is powered for effects of roughly
three percentage points and cannot exclude smaller ones;
the γ = 1.0 accuracy lean (positive in three of three seeds) is
unproven, not disproven. Second, held-out trajectories during training close the remaining escape route: linter compli-

## Page 5

6

Predictions Out of Sample

Section 4’s account was frozen, with quantitative bands, before a second pre-registered experiment: seven configurations probing the two injection paths the mechanism identifies, at fixed model, data, and optimization settings (fourteen
runs; configurations and predictions were committed to hashidentified manifests before any run started).

D1: removing the divisor restores dose response. The
mechanism predicts that disabling the standard-deviation
division (scale_rewards=none, retaining group meancentering) makes the process term’s advantage share β-linear
everywhere, within a factor of two of β/0.1 × 0.0344 (the
Regime-1 reference measured in Section 4). Measured (robust medians, per seed): 0.043/0.034 at β=0.1 against band
[0.017, 0.069]; 0.145/0.150 at β=0.5 against [0.086, 0.344];
0.206/0.237 at β=1.0 against [0.172, 0.688], with six of six
seeds in band and the γ=0 counterfactual at exactly zero in
both its seeds. The shares run 1:3.8:5.7 against β ratios of
1:5:10: β-monotone, but sub-linear, the largest cell falling
nearly 2× short of a strictly linear extrapolation. We read
this as self-damping (a large process term inflates the group
spread it is normalized against), but we flag that this is a posthoc reading, not part of the frozen prediction, which committed only to the factor-of-two bands; the bands are deliberately
loose and we report six cell-seeds (three β cells × two seeds,
not six independent draws) landing inside them, with the
caveat that a tighter linear band would have flagged the β=1.0
point. The companion prediction D2 also held: with the divisor gone, uniformly-wrong groups carry β-proportional advantages (0.048/0.127/0.208 median maxima) rather than
the saturated ±≈1 that standardization manufactured, and
exactly zero at γ=0, where outcome-only training gives such
groups no gradient at all.

median process-induced
standardized-advantage share

ance on a 200-problem held-out slice rose by +4.0pp in the
outcome-only arm (as much as in any process arm, +1.8 to
+3.5pp), and the correlation between γ and held-out linter
rise across the twelve runs is r = 0.04 (p = 0.89). Whatever
trace-formatting drift reinforcement learning induces here is
a property of optimizing for answers, not of the process term.
A reward-hacking axis evaluated per checkpoint (rising linter
compliance against flat held-out accuracy) triggered in none
of the twelve runs.
The null survives adversarial scrutiny by construction
rather than by assertion: a cross-model integrity audit (two
reviewer models from different families) verified groundtruth provenance, statistics, and the extraction pipeline, and
its findings (an extraction-completeness guarantee retrofitted
with hash-level verification, an overclaim in our own mechanism wording that we retracted, and denominator-artifact
hygiene in the diagnostics) are folded into every number
quoted here. Combined with Section 4, the result is a null
with a cause: the diagnostics show the process signal being
suppressed in one regime and saturated in the other, the endpoints show the consequence, and Section 6 tests whether the
same account predicts what happens when the suppression
is removed.

0.4

0.3

frozen 2 × band
cell-seed obs. (6/6 in band)
γ = 0 counterfactual

group-std. share, matched wt. (0.034)
strict-linear ref.

0.2

0.1

0.0

0.0 0.1
0.5
1.0
process weight β (at γ = 1)

Figure 1: D1: disabling the standardizing divisor restores
dose-responsive process influence. Process-term advantage
share (robust median over groups with σ R ≥ 0.05) versus the
process weight β at γ=1, under scale_rewards=none.
Points are six cell-seed observations (two seed identities);
shaded regions are the factor-of-two bands frozen before
the runs, and 6/6 observations fall inside them. The green
diamond is the γ=0 counterfactual (share exactly zero;
outcome-only training gives these groups no gradient).
The orange dashed line marks the share measured under
group standardization at the matched setting — the earlier
sweep’s maximum tested effective line-credit weight (γ=1.0,
γβ=0.1; robust median 0.034). The unwashed β=0.1 share
coincides with it and rises from there with β — weights
beyond this were unreachable in the standardized sweep,
whose bundle coefficient cancels in uniform-outcome groups
(Lemma 1). The rise is monotone but sub-linear against the
dotted strict-linear reference (the β=1 point falls short, a
post-hoc self-damping reading; see text). The x-axis is linear
so the counterfactual sits at its true x=0. Advantage-level prediction; endpoint improvement did not replicate (Section 6).

D4: a prediction that failed, reported as such. We preregistered an arbitration of an external result: a deterministic
step-verifier reward summed over many steps has been reported to improve a structured-reasoning task as an additive
composite under this same normalization family (Pronesti,
Belz, and Hou 2026), and our mechanism’s simplest reading — additive composites work when the aggregate process
share is large — predicted that our largest-share cell (β=1.0,
≈21–24% advantage share) would show the endpoint effect smaller-share cells lack. It did not (verified-correct rate
+0.7pp at the 256-token cap, p=0.13). The naive magnitude story is therefore incomplete: D1 and D2 confirm the
mechanism’s advantage-level predictions, but D4 was its one
endpoint prediction and it failed, so the mechanism is, at this
point, confirmed as an account of the advantage displacement
and unconfirmed as a predictor of endpoints. Two statuses
should be kept distinct here. The pre-registered operationalization of D4 failed by its frozen rule. As an arbitration of
the external result, however, the test is inconclusive rather

## Page 6

than negative: our β=1.0 cell differs from that setting in
group size, reward aggregation (trajectory-level versus perstep-summed), and domain, and we did not run their configuration at our scale (a direct probe we flag as the natural next
experiment). “Large aggregate share” is thus a hypothesis our
indirect test did not bear out, not an explanation we can assert
or refute. We keep the failed prediction in the record because
the confirmed ones are only as credible as our willingness to
report this one.

An endpoint candidate that did not survive its controls.
Restoring dose response produced one endpoint candidate.
The β=0.5 unwashed-additive configuration cleared our preregistered endpoint gate on the selection set (verified-correct
rate at the 256-token cap +1.90pp, Holm-adjusted p=0.006,
paired 95% CI [+0.8, +3.0]), and its mechanism signature
was in band (D1 above). Under the confirmation protocol
we committed to before running it, it did not survive, on
two independent grounds. First, on fresh seeds the effect
did not replicate (+0.00pp, McNemar b=c=19, p=1.0; fullgeneration accuracy leaned negative, −1.9pp, p=0.078): the
fresh-seed failure alone rejects the candidate. Second, the
within-group shuffled control, which preserves the process
reward’s magnitude distribution per group while destroying
its association with trace content, reproduced and slightly exceeded the candidate’s effect in every slice (+0.70pp vs. P0 at
the 256-token cap): whatever signal exists is not attributable
to what the verifier measures. By the binding rule fixed in
advance (ambiguity counts against the candidate), the candidate is rejected. We are careful about what the data do and do
not establish: they reject P2 and reject a semantic attribution
of its probe-stage gain, but they do not distinguish among
seed selection, optimization noise, and a magnitude effect as
the source of that gain, and we do not claim to. We report
the episode prominently because it is the sharpest evidence
for this paper’s thesis: a pre-registered, multiple-comparison-corrected, mechanism-consistent endpoint result did not survive fresh-seed replication and a semantic-association control. Measuring a verifiable-reasoning reward’s effect takes
more than a clean single-experiment gate; it takes replication and a control that separates semantics from magnitude. The mechanism results (D1/D2) are unaffected: they
are advantage-level claims, confirmed on their own preregistered bands and independent of any endpoint outcome.

D3: where token-level credit breaks. The second injection path bypasses the scalar entirely: token-level advantage offsets from the linter’s per-line verdicts, in our uncentered verifier-line token-offset formulation — tokens inside pass/fail/uncheckable lines receive a constant signed
offset, unnormalized and independent of the trajectory’s outcome. At low weight (γ tok =0.1) this is a mild, harmless
drift; at γ tok =0.5 it is destructive, and the damage is fully attributable. Three independent instruments agree. The tokenadvantage diagnostics show 50–60% of completion tokens
carrying offsets whose per-token distortion scales exactly
with dose (token-advantage deviation 0.050 → 0.245 for
a 5× weight increase, about half the scale of the groupnormalized scalar advantage, applied always-on). The emitted traces restructure accordingly: at γ tok =0.5 the policy

writes ∼22% fewer characters but ∼60% more equation
lines and ∼75% more check lines, halving its untagged
derivational prose. The reward has become an annuity on
checkable-line tokens, paying every token unconditionally,
where the outcome anchor pays once per trajectory, grouprelative. And held-out accuracy declines during training
in both seeds (−3.5 and −9.0pp first-to-last), so the harm
is an optimization pathology accruing over steps, not an
evaluation artifact; final accuracy lands 7.5pp below baseline. This is the hacking surface we pre-registered for this
formulation (“verbose-but-checkable padding”), realized as
checkable-line density rather than length, and caught inflight by the same instrumentation that measured the mechanism. We scope the finding precisely: it falsifies this uncentered, outcome-decoupled token-offset design at meaningful
weight, and marks centering and outcome-gating as the loadbearing choices any token-level verifier credit must get right
— not token-level credit as a class.

7

Implications and Limitations

For practitioners adding a process reward to GRPO, the
mechanism yields a concrete design rule: under per-group
standardization, a trajectory-additive process term cannot be
tuned into relevance by its coefficient — the coefficient cancels wherever the term is the only signal and is swamped
wherever the outcome varies. If the process signal is to matter at all, the standardization must be changed (removing the
divisor restores dose-response, as our D1 experiment confirms) or the signal injected off the scalar advantage entirely
— and our token-level attempt shows that the latter is not free:
an uncentered, outcome-decoupled per-token offset becomes
an annuity on checkable tokens and degrades the answers it
was meant to support. Centering and outcome-gating are the
load-bearing choices for any token-level verifier credit; we
did not find a setting of our formulation that helped.
The broader implication is methodological. Our surviving
endpoint candidate was pre-registered, corrected for multiple comparisons, significant at p=0.006, and consistent with
an independently confirmed mechanism — four properties
usually taken as strong evidence — and it still failed to replicate on fresh seeds and failed a magnitude-matched semantic
control. We argue the exposure is endemic, not incidental to
our setup. Any process reward adds reward mass correlated
with the process metric it optimizes; under group-relative
advantage estimation, adding mass to a subset of completions shifts their advantages whether or not the mass tracks
anything semantic, and the resulting endpoint movement is
observationally identical to a genuine effect on the metric the
reward and the evaluation share. This is a structural property
of shaped rewards under group normalization, not a quirk of
linter validity: any verifiable-reasoning reward whose target
metric correlates with its own magnitude inherits it. The standard defenses do not address it: a held-out set, multiple seeds
at train time, and multiple-comparison correction all leave
a magnitude-driven gain indistinguishable from a semantic
one, as our own p=0.006 candidate shows. The control that
separates the two is a within-group permutation that preserves the reward’s magnitude distribution while destroying
its association with trace content: if the permuted reward

## Page 7

reproduces the gain, the gain cannot be attributed to what
the verifier measures. We propose this shuffle as a standard,
near-zero-cost control for endpoint claims about verifiable
process rewards, and we release its implementation so it can
be dropped into an existing GRPO reward path.
Limitations bound these conclusions and we state them
without hedging. All experiments use a single 4B model
with QLoRA adapters on mathematical reasoning with a
compact trace language; the mechanism’s algebra is general
to group-standardized objectives, but its magnitudes — the
specific few-percent displacement share — are measured at
one scale, one group size, and one process-score scale, and
larger models or different verifiers may sit at different operating points. Our endpoint statistics are powered for effects of
roughly three percentage points and cannot exclude smaller
ones; we claim a null within that floor, not a proof of no
effect. The trace language restricts what the model may express, which bounds its accuracy ceiling independently of
the reward question; we report expressibility coverage in the
supplement but do not disentangle it here. And our arbitration of a positive external result with a deterministic step
verifier remains open: our largest-share configuration did not
reproduce that regime’s benefit, but the two settings differ in
group size, per-step versus trajectory-level reward aggregation, and domain, so we report the prediction as failed rather
than the external result as explained.

8

Reproducibility and Artifacts

Every experiment in this paper was pre-registered before it
ran: the sweep grid, the coefficient bands for the mechanism
predictions, the endpoint gate criterion with its multiplecomparison correction, and the confirmation protocol with
its binding shuffle-control rule were all committed to hashidentified configuration manifests and frozen plan documents, with dated records, prior to the corresponding runs.
The distinction between the selection set (on which all model
and coefficient choices were made) and the confirm-stage
analysis is preserved throughout, and the endpoint candidate’s death was decided by a rule fixed in advance of the
data.
The supplementary material contains the trace-language
grammar and the full rule specification of the linter; the reward and advantage-injection code for both fixes with the
unit tests that lock their invariants (including the coefficientzero equivalences and the within-group shuffle’s magnitudepreservation and group-locality); the per-cell diagnostic and
eval records with the analysis scripts that regenerate every number in the paper; the audit reports; and the preregistration documents with their revision history. All numbers reported here trace to those records through a zerocontext claim audit run before submission. The supplement
is self-contained (no external links) and anonymized.
We intend to publicly release the trace language, the linter, the advantage-injection implementations, and the instrumentation and analysis harness, so that both the mechanism
measurements and the confirmation protocol can be reproduced and reused as controls for verifiable-reasoning reward
claims.

References

Bay, Y. Y.; and Yearick, K. A. 2026. GRPO, Dr. GRPO, and
DAPO Are Three Operations on One Number: The Group-
Standard-Deviation Identity. arXiv:2607.00152.

Chen, H.; Yang, T.; Gao, S.; Chen, R.; Quan, X.; Tian, H.;
and Yao, T. 2025. Discriminative Policy Optimization for
Token-Level Reward Models. arXiv:2505.23363.

Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.;
Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.;
Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve
Math Word Problems. arXiv:2110.14168.

DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.
arXiv:2501.12948.

Ding, R.; Lv, Y.; Meng, X.; Song, J.; Wang, C.; Jiang, C.;
and Cheng, Y. 2026. PRPO: Aligning Process Reward with
Outcome Reward in Policy Optimization. arXiv:2601.07182.

Dumitru, R.-G.; Peteleaza, D.; Yadav, V.; and Pan, L. 2025.
ConciseRL: Conciseness-Guided Reinforcement Learning
for Efficient Reasoning Models. arXiv:2505.17250.

Gao, J.; Yan, S.; Tan, Q.; Yang, L.; Xu, S.; Fu, W.; Mei, Z.;
Lyu, K.; and Wu, Y. 2025. How Far Are We from Optimal
Reasoning Efficiency? arXiv:2506.07104.

Garg, A.; Zhang, C.; Neema, N.; Bick, D.; Venkatesh, G.;
and Hestness, J. 2025. CoRPO: Adding a Correctness Bias
to GRPO Improves Generalization. arXiv:2511.04439.

Ge, C.; Yin, C. H.; Liang, H.; and Zhang, J. 2026. Why
GRPO Needs Normalization: A Local-Curvature Perspective
on Adaptive Gradients. arXiv:2601.23135.

He, P.; Huang, Y.; Sachan, M.; and Jin, Z. 2026. Uncovering
Hidden Correctness in LLM Causal Reasoning via Symbolic
Verification. arXiv:2601.21210.

Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart,
S.; Tang, E.; Song, D.; and Steinhardt, J. 2021. Measuring
Mathematical Problem Solving With the MATH Dataset.
arXiv:2103.03874.

Le, T.-L. V.; Jeon, M.; Vu, K.; Lai, V.; and Yang, E. 2025.
No Prompt Left Behind: Exploiting Zero-Variance Prompts
in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping. arXiv:2509.21880.

Li, Z.; Kang, L.; Xiao, F.; Xing, L.; Si, Q.; Li, Z.; Gong, W.;
Yang, D.; Xiao, Y.; and Guo, H. 2026. Outcome-Grounded
Advantage Reshaping for Fine-Grained Credit Assignment
in Mathematical Reasoning. arXiv:2601.07408.

Liu, Z.; Chen, C.; Li, W.; Qi, P.; Pang, T.; Du, C.; Lee, W. S.;
and Lin, M. 2025. Understanding R1-Zero-Like Training: A
Critical Perspective. arXiv:2503.20783.

Parthasarathi, P.; Reymond, M.; Chen, B.; Cui, Y.; and Chandar, S. 2025. GRPO-λ: Credit Assignment improves LLM
Reasoning. arXiv:2510.00194.

Pronesti, M.; Belz, A.; and Hou, Y. 2026. Beyond Outcome
Verification: Verifiable Process Reward Models for Structured Reasoning. arXiv:2601.17223.

## Page 8

Salmani-Zarchi, M. M.; Rahimi, Z.; Faili, H.; and Dousti,
M. 2026. MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following.
arXiv:2606.06058.
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.-M.; Zhang, M.;
Li, Y. K.; Wu, Y.; and Guo, D. 2024. DeepSeekMath: Pushing
the Limits of Mathematical Reasoning in Open Language
Models. arXiv:2402.03300.
Su, C. 2025. MedRule-KG: A Knowledge-Graph-Steered
Scaffold for Mathematical Reasoning with a Lightweight
Verifier. arXiv:2510.16309.
Sui, Y.; Chuang, Y.-N.; Wang, G.; Zhang, J.; Zhang, T.;
Yuan, J.; Liu, H.; Wen, A.; Zhong, S.; Chen, H.; and Hu, X.
2025. Stop Overthinking: A Survey on Efficient Reasoning
for Large Language Models. arXiv:2503.16419.
Wang, C.; Tian, H.; Yang, T.; Shi, Y.; Yao, T.; and Ding,
W. 2026a. Process Advantage Signal Shaping: A Paradigm-
Agnostic Middleware for Process-Supervised RL in LLM
Reasoners. arXiv:2606.29296.
Wang, S.; Pham, Q. H.; Yin, F.; Wang, X.; Chen,
J. Q.; Durrett, G.; and Ye, X. 2026b. Detecting and
Suppressing Reward Hacking with Gradient Fingerprints.
arXiv:2604.16242.
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Chi, E. H.;
Xia, F.; Le, Q.; and Zhou, D. 2022. Chain of Thought
Prompting Elicits Reasoning in Large Language Models.
arXiv:2201.11903.
Wen, H.; Wu, X.; Sun, Y.; Zhang, F.; Chen, L.; Wang, J.; Liu,
Y.; Liu, Y.; Zhang, Y.-Q.; and Li, Y. 2025. BudgetThinker:
Empowering Budget-aware LLM Reasoning with Control
Tokens. arXiv:2508.17196.
Xie, G.; Shi, Y.; Tian, H.; Yao, T.; and Zhang, X. 2025.
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment. arXiv:2508.02298.
Yi, J.; and Wang, J. 2025. ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient
Reasoning. arXiv:2504.21370.
Yu, Q.; Zhang, Z.; Zhu, R.; Yuan, Y.; Zuo, X.; Yue, Y.;
et al. 2025. DAPO: An Open-Source LLM Reinforcement
Learning System at Scale. arXiv:2503.14476.
Zhang, Q.; Guo, T.; Ren, X.; Chen, J.; Ding, M.; Xin, R.; and
Xiao, X. 2026. Scaling Reasoning Tokens via RL and Parallel Thinking: Evidence From Competitive Programming.
arXiv:2604.01302.
Zhou, H.; Yang, A. X.; Aitchison, L.; Korhonen, A.; and
Jiang, A. Q. 2026. Reasoning Arena: Trace Tournaments
When Verifiable Rewards Fall Short. arXiv:2606.09380.
