{
  "slug": "additive-process-rewards",
  "title": "Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It",
  "short": "APR",
  "line": "reasoning",
  "line_name": "Verifiable reasoning",
  "part": null,
  "status": "AAAI 2027 submission",
  "date": "2026-07-21",
  "authors": [
    "Sisong Bei",
    "Mikhail L Arbuzov",
    "Ziwei Dong",
    "Dmitri Kalaev",
    "Alexey Shvets"
  ],
  "abstract": "Verifiable process rewards—dense signals from checking intermediate reasoning steps—are widely expected to improve reinforcement learning of language-model reasoning. We show that in the standard group-normalized policyoptimization pipeline (GRPO), trajectory-additive process rewards are suppressed, in two regimes. In groups whose outcomes agree, the standardizing coefficient cancels exactly— an analytic consequence of the group-standard-deviation algebra—making the process term’s magnitude un-tunable, so coefficient sweeps are binary by construction. In groups whose outcomes disagree, the process term is compressed to a few percent of the advantage—empirically, a median of 0.3– 6% for process scores whose within-group spread is small relative to the outcome spread, as in our setting. Measuring this requires a verifier fast and deterministic enough to instrument every reward call: we use a compact machine-checkable trace language whose rule-based linter catches 99.5% of injected rule-family reasoning errors (vs. 66.3% for self-verification and 87.2% for a frontier LLM judge) at 9 ms per trace with no model in the loop—an operating point complementary to proof assistants, which verify formalized proofs rather than native traces. In a pre-registered 12-run sweep (4B model, mathematical reasoning), the mechanism’s predictions hold and endpoint effects are null within a ∼3pp detection floor. We derive and test two pre-registered fixes and an external-result arbitration; a separately pre-registered endpoint candidate that cleared multiple-comparison correction did not survive fresh-seed replication and a magnitude-matched semantic control.",
  "tldr": "Under GRPO's per-group standardization, a process term added to the outcome reward cannot be tuned: its weight cancels where a group's outcomes agree and is swamped where they differ. Endpoint gains need replication and a shuffle control.",
  "pages": 8,
  "html": "https://telegrapher.ai/research/additive-process-rewards/",
  "md": "https://telegrapher.ai/research/additive-process-rewards.md",
  "reader": "https://telegrapher.ai/research/additive-process-rewards/read/",
  "pdf": "https://telegrapher.ai/papers/additive-process-rewards/additive-process-rewards.pdf",
  "arxiv": null,
  "openreview": null,
  "post": "https://telegrapher.ai/blog/additive-process-rewards/",
  "bibtex": "@misc{bei2026why,\n  title         = {Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It},\n  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},\n  year          = {2026},\n  note          = {AAAI 2027 submission},\n  url           = {https://telegrapher.ai/research/additive-process-rewards/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "In GRPO, an additive process reward's weight cancels if outcomes agree and is swamped if not"
    },
    {
      "label": "TL;DR",
      "text": "Under GRPO's per-group standardization, a process term added to the outcome reward cannot be tuned: its weight cancels where a group's outcomes agree and is swamped where they differ. Endpoint gains need replication and a shuffle control."
    },
    {
      "label": "Method",
      "text": "Twelve GRPO runs of Qwen3.5-4B with QLoRA adapters (process weight γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each, group size G=4, 1,500 steps) were rewarded with a binary outcome plus a linter-based process score, evaluated greedily on the same 500 MATH-500 problems, and followed by a pre-registered fourteen-run probe of two fixes whose endpoint candidate went through a frozen fresh-seed and within-group-shuffle confirmation."
    },
    {
      "label": "Key result",
      "text": "Process-term displacement of standardized advantages in mixed-outcome groups: 0.3–6%; Training groups with zero advantage at γ = 0: 47.8%; Linter detection rate on injected errors: 99.5%"
    },
    {
      "label": "Why it matters",
      "text": "Anyone adding a verifier score to a GRPO reward as an additive term gets a short design rule."
    },
    {
      "label": "Limits",
      "text": "Every run uses one model, Qwen3.5-4B with QLoRA adapters, on mathematical reasoning written in one compact trace language."
    },
    {
      "label": "Status",
      "text": "AAAI 2027 submission, July 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/additive-process-rewards/read/, PDF /papers/additive-process-rewards/additive-process-rewards.pdf"
    }
  ],
  "note": {
    "slug": "additive-process-rewards",
    "claim_title": "In GRPO, an additive process reward's weight cancels if outcomes agree and is swamped if not",
    "meta_description": "Under GRPO's group standardization, an additive process reward's coefficient cancels in uniform-outcome groups and shrinks to a few percent in mixed ones.",
    "tldr": "Under GRPO's per-group standardization, a process term added to the outcome reward cannot be tuned: its weight cancels where a group's outcomes agree and is swamped where they differ. Endpoint gains need replication and a shuffle control.",
    "gist": "Adding a verifier's step-level score to the outcome reward looks like a natural way to reward derivations that check out. Under GRPO's per-group reward standardization, such a trajectory-additive process term lands in one of two regimes. When every completion in a group gets the same outcome, the process score becomes the whole advantage, yet its coefficient cancels, so any positive weight yields the same advantages. When outcomes differ, their spread sets the denominator and the process term moves advantages by a median of 0.3–6%. A coefficient sweep therefore behaves like an on/off treatment. Measuring this took a rule-based linter that checks a compact trace language in 9 ms per trace with no model in the loop. In a pre-registered 12-run sweep of a 4B model on mathematical reasoning, accuracy showed no effect within a detection floor of about 3pp. Removing the standard-deviation divisor restored an advantage-level dose response; the one endpoint gain that followed failed fresh-seed replication and a magnitude-preserving shuffle control. All of it comes from one model, one group size and one domain.",
    "method": "Twelve GRPO runs of Qwen3.5-4B with QLoRA adapters (process weight γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each, group size G=4, 1,500 steps) were rewarded with a binary outcome plus a linter-based process score, evaluated greedily on the same 500 MATH-500 problems, and followed by a pre-registered fourteen-run probe of two fixes whose endpoint candidate went through a frozen fresh-seed and within-group-shuffle confirmation.",
    "summary_html": [
      "Take a group of four completions for one math problem, all of them wrong. GRPO subtracts the group's mean reward and divides by its standard deviation, so an outcome reward that is zero for every completion drops out of both, and the advantages come from the process score alone, divided by its own spread. Raise the process weight tenfold and numerator and denominator rise together. That case, and its opposite, where outcomes split and their spread sets the denominator, are the two regimes measured here for a <em>trajectory-additive</em> reward <code>R = r_ans + γs</code>. The process score <code>s</code> came from a rule-based linter for a compact trace language of givens, goals, tagged equations and explicit checks. It verifies a trace in 9 ms at the median, with a computer-algebra call and no model in the loop, which is fast enough to log group statistics at every reward call. Twelve Qwen3.5-4B policies were trained at γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each, and evaluated on the same 500 MATH-500 problems; a second pre-registered experiment of fourteen runs then tested what the account predicted.",
      "In mixed-outcome groups the process term moved standardized advantages by a median of 0.3–6% across γ, while the outcome moved them by about one unit. In uniform-outcome groups γ cancels exactly, a corollary of the shared-denominator algebra that the paper does not claim as new. The two regimes interact. At γ = 0, 47.8% of training groups had zero advantage; any positive γ switched on those whose process scores differed, and the zero-advantage share fell to ≈32% at γ = 0.1, then stayed put through γ = 1.0. The sweep was close to a single on/off treatment, and its endpoints were null: the largest accuracy contrast, +1.47pp at γ = 1.0, had a 95% interval crossing zero. Dropping the standard-deviation divisor brought back a dose response at the advantage level, monotone but sub-linear, with six of six cell-seed observations inside the frozen bands. It also produced one endpoint candidate, +1.90pp verified-correct at Holm-adjusted p=0.006. On fresh seeds that gain was +0.00pp, and a within-group shuffled reward with the same per-group magnitudes reproduced and slightly exceeded the candidate's effect in every slice. A token-level alternative, at weight 0.5, padded traces with checkable lines and finished 7.5pp below baseline accuracy."
    ],
    "key_numbers": [
      {
        "label": "Process-term displacement of standardized advantages in mixed-outcome groups",
        "value": "0.3–6%",
        "context": "median across γ, groups with σ_R ≥ 0.05; outcome-driven displacements are on the order of one unit",
        "bad": true
      },
      {
        "label": "Training groups with zero advantage at γ = 0",
        "value": "47.8%",
        "context": "18.0% of all groups carry process scores that could separate them; any γ > 0 cuts the zero-advantage share to ≈32%, unchanged up to γ = 1.0"
      },
      {
        "label": "Linter detection rate on injected errors",
        "value": "99.5%",
        "context": "at a 10% false-positive rate, 9 ms per trace; self-verification 66.3%, frontier LLM judge 87.2%; rates not false-positive-matched and errors drawn from the linter's own rule vocabulary"
      },
      {
        "label": "Largest accuracy contrast in the 12-run sweep (γ = 1.0 vs γ = 0)",
        "value": "+1.47pp",
        "context": "paired 95% CI [−0.27, +3.20], McNemar p = 0.115 over 1,500 paired problem-seed observations (500 MATH-500 problems, three seeds); design powered for roughly 3pp"
      },
      {
        "label": "Endpoint candidate's verified-correct gain on the selection set",
        "value": "+1.90pp",
        "context": "Holm-adjusted p=0.006 at the 256-token cap; +0.00pp on fresh seeds, and a magnitude-preserving shuffled reward reproduced and slightly exceeded the candidate's effect in every slice",
        "bad": true
      }
    ],
    "editorial_html": [
      "Anyone adding a verifier score to a GRPO reward as an additive term gets a short design rule. Under per-group standardization the coefficient is not a dial: it cancels wherever the process term is the only signal and is swamped wherever the outcome varies, so a sweep mostly tests whether the weight is above zero. A flat sweep is a null with a cause, not a verdict on process signals. For the signal to count at all, the standardization has to change (dropping the divisor restored dose response at the advantage level here) or the credit has to enter off the scalar advantage. That second route is not free. An uncentered per-token offset, paid whatever the outcome, became an annuity on checkable tokens and degraded the answers it was meant to support; centering and outcome-gating are the load-bearing choices.",
      "The larger point is about evidence. The endpoint candidate was pre-registered, corrected for multiple comparisons, significant at p=0.006 and consistent with an independently confirmed mechanism. It still failed. The paper argues the exposure is general: any process reward adds reward mass that correlates with the metric it targets, and under group-relative advantages that mass shifts advantages whether or not it tracks anything semantic; held-out sets, extra training seeds and multiple-comparison correction all leave such a gain looking semantic. A within-group permutation of the process reward keeps its magnitude distribution and breaks its link to trace content, which separates the two. The paper proposes it as a standard, near-zero-cost control for endpoint claims about verifiable process rewards.",
      "The trace language and its linter come from <em>Telegraph Reasoning</em>, which introduces a grammar that lets a rule-based linter and a symbolic algebra system check every step of a trace; here they are the measuring instrument rather than the result. The same language restricts what a model may write, and so bounds accuracy apart from any reward. <em>Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning</em> takes up that cost, scoring checker acceptance and answer accuracy on the same outputs against free-form reasoning, and finds that the two come apart."
    ],
    "limitations": "Every run uses one model, Qwen3.5-4B with QLoRA adapters, on mathematical reasoning written in one compact trace language. The cancellation in uniform-outcome groups follows from the algebra of any group-standardized objective, but the few-percent displacement in mixed groups is an empirical magnitude: it depends on process scores varying much less within a group than outcomes do, and it was measured at one scale, one group size (G=4) and one process-score scale, so larger models or other verifiers may sit elsewhere. The endpoint statistics are powered for effects of roughly three percentage points. The sweep is a null within that floor, and the γ = 1.0 accuracy lean, positive in three of three seeds, is unproven rather than disproven; the pooled problem-seed test also treats each checkpoint as fixed, so its intervals are narrower than a seed-random-effects analysis would give. The linter's detection comparison is not false-positive-matched and draws its errors from the linter's own rule vocabulary. The trace language bounds accuracy independently of the reward, and the paper does not disentangle the two. Reading the sub-linear dose response as self-damping is post hoc. The arbitration of an external positive result with a deterministic step verifier is inconclusive: the largest-share configuration did not reproduce that benefit, but group size, reward aggregation and domain differ, and the external configuration was not run at this scale. And the shuffle control rejects a semantic reading of the candidate's gain without saying whether seed selection, optimization noise or a magnitude effect produced it.",
    "concept_terms": [
      "group-relative policy optimization (GRPO)",
      "advantage standardization",
      "verifiable process rewards",
      "zero-variance groups",
      "within-group shuffle control",
      "pre-registered replication"
    ],
    "references": [
      "Shao et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.",
      "Liu et al. (2025). Understanding R1-Zero-Like Training: A Critical Perspective.",
      "Bay and Yearick (2026). GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity.",
      "Pronesti et al. (2026). Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning.",
      "Wang et al. (2026). Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/additive-process-rewards.txt",
    "verification": {
      "claims": [
        {
          "claim": "GRPO converts rewards to advantages by standardizing within each sampled group",
          "verdict": "CONFIRMED",
          "source_quote": "rewards to advantages by standardizing within the group"
        },
        {
          "claim": "The reward studied is a trajectory-additive composite R = r_ans + γs with a binary outcome term",
          "verdict": "CONFIRMED",
          "source_quote": "where r ans ∈ {0, 1} is the verifiable outcome"
        },
        {
          "claim": "In uniform-outcome groups the advantage reduces to γ(s_i − s̄)/(γσ_s + ϵ), so raising γ raises numerator and denominator together",
          "verdict": "CONFIRMED",
          "source_quote": "A i = γ(s i − s̄)/(γσ s + ϵ)"
        },
        {
          "claim": "In uniform-outcome groups the process term is the whole standardized advantage",
          "verdict": "CONFIRMED",
          "source_quote": "it is the entire standardized advantage, at full magnitude"
        },
        {
          "claim": "In uniform-outcome groups any positive γ yields the same advantages (coefficient cancels)",
          "verdict": "CONFIRMED",
          "source_quote": "any γ > 0 produces the same advantages as any other"
        },
        {
          "claim": "The coefficient cancels exactly (paper's wording; Lemma 1 states it as the γσ_s/ϵ → ∞ limit)",
          "verdict": "CONFIRMED",
          "source_quote": "the coefficient cancels exactly"
        },
        {
          "claim": "Lemma 1 is presented as a corollary of the shared-denominator algebra, not as a new result",
          "verdict": "CONFIRMED",
          "source_quote": "we present it as such rather than as a novelty"
        },
        {
          "claim": "Mixed-outcome groups: process term displaces standardized advantages by a median of 0.3–6% across γ (σ_R ≥ 0.05)",
          "verdict": "CONFIRMED",
          "source_quote": "term displaces post-normalization advantages by a median of 0.3–6% across γ"
        },
        {
          "claim": "Outcome-driven displacements are on the order of one unit",
          "verdict": "CONFIRMED",
          "source_quote": "outcome-driven displacements on the order of one unit"
        },
        {
          "claim": "The few-percent magnitude is empirical, conditional on small within-group process spread, measured at one scale and G=4",
          "verdict": "CONFIRMED",
          "source_quote": "measured at one scale, group size (G=4), and model class"
        },
        {
          "claim": "At γ = 0, 47.8% of training groups had zero advantage",
          "verdict": "CONFIRMED",
          "source_quote": "At γ = 0, 47.8% of training groups produce identically zero advantage"
        },
        {
          "claim": "18.0% of all groups carry process scores that could separate their completions",
          "verdict": "CONFIRMED",
          "source_quote": "18.0% of all groups"
        },
        {
          "claim": "Any positive γ activates those groups; zero-advantage share falls to ≈32% at γ = 0.1",
          "verdict": "CONFIRMED",
          "source_quote": "the zero-advantage fraction drops to ≈32% at γ = 0.1"
        },
        {
          "claim": "Zero-advantage share stays put through γ = 1.0",
          "verdict": "CONFIRMED",
          "source_quote": "does not move further as γ grows to 1.0"
        },
        {
          "claim": "A coefficient sweep behaves as an approximately binary (on/off) treatment",
          "verdict": "CONFIRMED",
          "source_quote": "behaves as an approximately binary treatment"
        },
        {
          "claim": "Trace language of givens, goals, tagged equations and explicit checks",
          "verdict": "CONFIRMED",
          "source_quote": "the model states givens, goals, tagged equations, and explicit checks"
        },
        {
          "claim": "Linter is pure rules plus a computer-algebra call, no model in the loop",
          "verdict": "CONFIRMED",
          "source_quote": "pure rules plus a computer-algebra call"
        },
        {
          "claim": "Linter verifies a trace in 9 ms at the median",
          "verdict": "CONFIRMED",
          "source_quote": "verifies a trace in 9 ms at the median"
        },
        {
          "claim": "Group statistics are logged at every reward call",
          "verdict": "CONFIRMED",
          "source_quote": "These within-group statistics are logged at every reward call"
        },
        {
          "claim": "Linter detects 99.5% of injected errors at a 10% false-positive rate",
          "verdict": "CONFIRMED",
          "source_quote": "the linter detects 99.5% of errors at a 10% false-positive rate"
        },
        {
          "claim": "Self-verification 66.3%, frontier LLM judge 87.2%",
          "verdict": "CONFIRMED",
          "source_quote": "66.3% for self-verification and 87.2% for a frontier LLM judge"
        },
        {
          "claim": "Detection rates are not false-positive-matched",
          "verdict": "CONFIRMED",
          "source_quote": "rates that are not false-positive-matched across the three verifiers"
        },
        {
          "claim": "Injected errors are drawn from the linter's own rule vocabulary",
          "verdict": "CONFIRMED",
          "source_quote": "the injected errors are drawn from the same failure vocabulary the linter’s rules encode"
        },
        {
          "claim": "Twelve policies at γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each",
          "verdict": "CONFIRMED",
          "source_quote": "We trained twelve policies—γ ∈ {0, 0.1, 0.5, 1.0}, three seeds each"
        },
        {
          "claim": "Qwen3.5-4B with QLoRA adapters, group size G=4, 1,500 steps",
          "verdict": "CONFIRMED",
          "source_quote": "(Qwen3.5-4B) with QLoRA adapters, group size G=4, 1,500"
        },
        {
          "claim": "All twelve cells evaluated greedily on the same 500 MATH-500 problems",
          "verdict": "CONFIRMED",
          "source_quote": "evaluated greedily on the same 500 MATH-500 problems"
        },
        {
          "claim": "Pre-registered 12-run sweep of a 4B model; endpoints null within a ~3pp detection floor",
          "verdict": "CONFIRMED",
          "source_quote": "endpoint effects are null within a ∼3pp detection floor"
        },
        {
          "claim": "Largest accuracy contrast +1.47pp at γ = 1.0 vs γ = 0, paired 95% CI [−0.27, +3.20]",
          "verdict": "CONFIRMED",
          "source_quote": "is +1.47pp with a paired 95% confidence interval of [−0.27, +3.20]"
        },
        {
          "claim": "McNemar p = 0.115 over 1,500 paired problem-seed observations",
          "verdict": "CONFIRMED",
          "source_quote": "McNemar p = 0.115 over 1,500 paired problem-seed observations"
        },
        {
          "claim": "Design powered for effects of roughly three percentage points",
          "verdict": "CONFIRMED",
          "source_quote": "the design is powered for effects of roughly three percentage points"
        },
        {
          "claim": "γ = 1.0 accuracy lean positive in three of three seeds; unproven, not disproven",
          "verdict": "CONFIRMED",
          "source_quote": "positive in three of three seeds) is unproven, not disproven"
        },
        {
          "claim": "Pooled problem-seed test treats checkpoints as fixed; intervals narrower than a seed-random-effects analysis",
          "verdict": "CONFIRMED",
          "source_quote": "so its intervals are narrower than a seed-random-effects analysis would give"
        },
        {
          "claim": "A second pre-registered experiment of fourteen runs tested the account's predictions",
          "verdict": "CONFIRMED",
          "source_quote": "(fourteen runs; configurations and predictions were committed"
        },
        {
          "claim": "Two pre-registered fixes were probed",
          "verdict": "CONFIRMED",
          "source_quote": "We derive and test two pre-registered fixes"
        },
        {
          "claim": "Dropping the standard-deviation divisor restored an advantage-level dose response",
          "verdict": "CONFIRMED",
          "source_quote": "disabling the standardizing divisor restores dose-responsive process influence"
        },
        {
          "claim": "The restored dose response is monotone but sub-linear",
          "verdict": "CONFIRMED",
          "source_quote": "β-monotone, but sub-linear"
        },
        {
          "claim": "Six of six cell-seed observations inside the frozen bands",
          "verdict": "CONFIRMED",
          "source_quote": "with six of six seeds in band"
        },
        {
          "claim": "Reading the sub-linear response as self-damping is post hoc",
          "verdict": "CONFIRMED",
          "source_quote": "we flag that this is a posthoc reading"
        },
        {
          "claim": "Endpoint candidate: +1.90pp verified-correct at the 256-token cap, Holm-adjusted p=0.006, on the selection set",
          "verdict": "CONFIRMED",
          "source_quote": "rate at the 256-token cap +1.90pp, Holm-adjusted p=0.006"
        },
        {
          "claim": "On fresh seeds the candidate's gain was +0.00pp",
          "verdict": "CONFIRMED",
          "source_quote": "on fresh seeds the effect did not replicate (+0.00pp"
        },
        {
          "claim": "A within-group shuffled reward preserving per-group magnitudes reproduced and slightly exceeded the candidate's effect in every slice",
          "verdict": "CONFIRMED",
          "source_quote": "reproduced and slightly exceeded the candidate’s effect in every slice"
        },
        {
          "claim": "The shuffle preserves magnitude distribution per group while destroying association with trace content",
          "verdict": "CONFIRMED",
          "source_quote": "preserves the process reward’s magnitude distribution per group while destroying its association with trace content"
        },
        {
          "claim": "Confirmation protocol fixed before running it",
          "verdict": "CONFIRMED",
          "source_quote": "Under the confirmation protocol we committed to before running it"
        },
        {
          "claim": "Data do not distinguish among seed selection, optimization noise and a magnitude effect",
          "verdict": "CONFIRMED",
          "source_quote": "seed selection, optimization noise, and a magnitude effect"
        },
        {
          "claim": "Token-level alternative at weight 0.5 finished 7.5pp below baseline accuracy",
          "verdict": "CONFIRMED",
          "source_quote": "final accuracy lands 7.5pp below baseline"
        },
        {
          "claim": "Token-level alternative padded traces with checkable lines",
          "verdict": "CONFIRMED",
          "source_quote": "realized as checkable-line density rather than length"
        },
        {
          "claim": "Uncentered per-token offset became an annuity on checkable tokens and degraded answers",
          "verdict": "CONFIRMED",
          "source_quote": "an annuity on checkable tokens and degrades the answers it was meant to support"
        },
        {
          "claim": "Centering and outcome-gating are the load-bearing choices",
          "verdict": "CONFIRMED",
          "source_quote": "Centering and outcome-gating are the load-bearing choices"
        },
        {
          "claim": "Coefficient cancels where the process term is the only signal and is swamped where the outcome varies",
          "verdict": "CONFIRMED",
          "source_quote": "wherever the term is the only signal and is swamped"
        },
        {
          "claim": "Process signal matters only if the standardization changes or credit enters off the scalar advantage",
          "verdict": "CONFIRMED",
          "source_quote": "the standardization must be changed"
        },
        {
          "claim": "A flat sweep is a null with a cause",
          "verdict": "CONFIRMED",
          "source_quote": "the result is a null with a cause"
        },
        {
          "claim": "Candidate was pre-registered, corrected, significant at p=0.006, consistent with an independently confirmed mechanism, and still failed",
          "verdict": "CONFIRMED",
          "source_quote": "significant at p=0.006, and consistent with an independently confirmed mechanism"
        },
        {
          "claim": "Paper argues any process reward adds reward mass correlated with its target metric (argued, not measured)",
          "verdict": "CONFIRMED",
          "source_quote": "We argue the exposure is endemic"
        },
        {
          "claim": "Held-out sets, extra seeds and multiple-comparison correction leave a magnitude-driven gain indistinguishable from a semantic one",
          "verdict": "CONFIRMED",
          "source_quote": "leave a magnitude-driven gain indistinguishable from a semantic one"
        },
        {
          "claim": "Paper proposes the within-group shuffle as a standard, near-zero-cost control",
          "verdict": "CONFIRMED",
          "source_quote": "We propose this shuffle as a standard, near-zero-cost control"
        },
        {
          "claim": "Endpoint gains need replication and a shuffle control",
          "verdict": "CONFIRMED",
          "source_quote": "it takes replication and a control that separates semantics from magnitude"
        },
        {
          "claim": "The trace language used is Telegraph Reasoning",
          "verdict": "CONFIRMED",
          "source_quote": "We use Telegraph Reasoning (TE), a compact trace language"
        },
        {
          "claim": "Telegraph Reasoning introduces a grammar that lets a rule-based linter and a symbolic algebra system check every step (other paper; abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "[abstracts.yaml] a discrete grammar for reasoning traces that lets a small rule-based linter and a symbolic algebra system check every step"
        },
        {
          "claim": "Trace language bounds accuracy independently of the reward; not disentangled",
          "verdict": "CONFIRMED",
          "source_quote": "which bounds its accuracy ceiling independently of the reward question"
        },
        {
          "claim": "Measuring Answer Accuracy and Trace Verifiability scores acceptance and accuracy on the same outputs against free-form reasoning (other paper; abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "[abstracts.yaml] We measured answer accuracy and checker acceptance over the same outputs"
        },
        {
          "claim": "Measuring Answer Accuracy and Trace Verifiability finds checkability and correctness come apart (other paper; abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "[abstracts.yaml] checkability and correctness come apart"
        },
        {
          "claim": "All runs use one 4B model with QLoRA adapters on mathematical reasoning",
          "verdict": "CONFIRMED",
          "source_quote": "All experiments use a single 4B model with QLoRA adapters"
        },
        {
          "claim": "Cancellation is general to group-standardized objectives",
          "verdict": "CONFIRMED",
          "source_quote": "the mechanism’s algebra is general to group-standardized objectives"
        },
        {
          "claim": "Magnitudes measured at one scale, one group size, one process-score scale",
          "verdict": "CONFIRMED",
          "source_quote": "one scale, one group size, and one process-score scale"
        },
        {
          "claim": "External-result arbitration is inconclusive rather than negative",
          "verdict": "CONFIRMED",
          "source_quote": "the test is inconclusive rather than negative"
        },
        {
          "claim": "Largest-share configuration did not reproduce the external benefit; settings differ in group size, aggregation, domain",
          "verdict": "CONFIRMED",
          "source_quote": "our largest-share configuration did not reproduce that regime’s benefit"
        },
        {
          "claim": "External configuration not run at this scale",
          "verdict": "CONFIRMED",
          "source_quote": "we did not run their configuration at our scale"
        },
        {
          "claim": "Reference: Shao et al. (2024) DeepSeekMath",
          "verdict": "CONFIRMED",
          "source_quote": "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models"
        },
        {
          "claim": "Reference: Liu et al. (2025) Understanding R1-Zero-Like Training",
          "verdict": "CONFIRMED",
          "source_quote": "Understanding R1-Zero-Like Training: A Critical Perspective"
        },
        {
          "claim": "Reference: Bay and Yearick (2026) Group-Standard-Deviation Identity",
          "verdict": "CONFIRMED",
          "source_quote": "DAPO Are Three Operations on One Number: The Group- Standard-Deviation Identity"
        },
        {
          "claim": "Reference: Pronesti et al. (2026) Beyond Outcome Verification",
          "verdict": "CONFIRMED",
          "source_quote": "Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning"
        },
        {
          "claim": "Reference: Wang et al. (2026a) Process Advantage Signal Shaping",
          "verdict": "CONFIRMED",
          "source_quote": "Process Advantage Signal Shaping: A Paradigm- Agnostic Middleware for Process-Supervised RL in LLM Reasoners"
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}