telegrapher

Math traces that passed a deterministic checker had the right answer about one time in three

Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

ICLR 2027 submission, August 2026

On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.

What we did and found

One trace spends 29 prose steps hunting for the value of a sum, tries 80? 81? and settles on 80. Then it writes EQ[eq_a]: S = 80 and CHECK[c1]: arith: eq_a.value == 80. Both claims check, the revised checker returns VERIFIED, and the grader marks the answer wrong, because the check confirms agreement with a value the trace asserted, not how that value follows from the givens. The trace language makes part of a model's output executable: GIVEN binds variables, GOAL names the target, STEP carries prose, EQ and CHECK expose claims a program evaluates, and ANS gives the answer. A deterministic checker's accepting verdict is called VERIFIED, and its share of all outputs the verified fraction. A separate grader takes another route through the same output: it extracts the final answer and compares it with the reference. Both were run over Qwen3.5-4B outputs in two studies. A panel of 25 trained checkpoints at 2,048 tokens, scored with the original v1 checker, supplies the joint distribution. Six seed pairs train the trace interface and free-form chain-of-thought from the same base model, data order and outcome-only reward, with no reward for acceptance.

On the panel's 17,025 pooled outputs, acceptance picked out a more accurate subset: 35.0% of accepted outputs were correct, against 18.38% of all outputs. A real filter, and a leaky one. Wrong-and-accepted outputs outnumbered correct-and-accepted ones, and the rejected column held 2,446 correct answers, so the filter kept 21.8% of the correct answers available. Adopting the trace interface changed the pool itself. Matched on training examples, it raised verified fraction by 34.1 points and lowered held-out accuracy by 15.6; matched on training tokens, the gaps were 35.1 and 19.1, and all six pairs shared both signs under both matches. Since free-form output does not target the trace language, its acceptance is zero and the acceptance gain equals the trace arm's own rate. Both gaps were already present at the first measured checkpoint. Within every trace run accuracy ended higher than it started, while acceptance moved up, down or not at all. Across checkpoints, the registered test for a negative accuracy–acceptance association did not confirm: the rank correlation was −0.087 against a target of −0.40 or lower, and none of 12 checkpoint pairs matched on accuracy and output length differed by the required five points of verified fraction.

Key numbers

Accuracy among checker-accepted outputsv1 checker, 17,025 pooled checkpoint–problem outputs at 2,048 tokens; 18.38% across all outputs35.0%
Share of correct answers whose trace was acceptedsame panel; the rejected outputs held 2,446 correct answers21.8%
Held-out accuracy change from adopting the trace interfacemean of six seed pairs at matched training examples, 200 held-out problems; −19.1 pp at matched training tokens−15.6 pp
Verified-fraction change from adopting the trace interfacesame six pairs and match; free-form acceptance is zero, so this is the trace arm's own rate; +35.1 pp at matched training tokens+34.1 pp
Rank correlation of checkpoint accuracy with acceptance25 checkpoints; 95% interval −0.4576 to +0.3091; the registered target was −0.40 or lower with an interval excluding zero−0.087

What this does not show

The trained-model evidence covers one model, Qwen3.5-4B, on math, and acceptance means passing one checker's supported grammar and arithmetic rules; a checked claim is a specified test on a declared statement, not a check that the answer follows from the givens. The interface comparison changes the trace language, prompting, decoding constraints and answer channel together, so it cannot say which component costs accuracy. Its first measurement comes at step 100, after training had begun. The study grew from three seed pairs to six after the first three results were visible, the paired means carry no population-level interval, and the trajectories are repeated readings of those six pairs rather than further replicates. The paired export does not record its checker version, and missing raw paired outputs and trainer state limit replay. The checkpoint test deviated from its registration too: it resampled checkpoints that share training lineages instead of problems, and matched pairs greedily rather than optimally.

Every number above was checked against the paper text.
The authors' abstract

Reasoning traces expose intermediate statements, but accepting a trace and obtaining a correct answer are distinct evaluation outcomes. We built a deterministic checker for a small model’s math reasoning traces and found that checkability and correctness come apart: a trace the checker accepts is correct only about one time in three. The model was trained by outcome-only reinforcement learning to write a compact trace language whose equations and checks a program evaluates. We measured answer accuracy and checker acceptance over the same outputs, using a checkpoint panel and seed-paired runs against free-form reasoning at matched token budgets. At matched training-example exposure, the trace interface raises acceptance by about 34 points and lowers accuracy by about 16 points relative to free-form reasoning. The directions of these differences are already present at the first measured checkpoint. Across the model’s own checkpoints, the accuracy–checkability trade-off we pre-registered did not confirm. Checkability and correctness are different quantities and should be measured separately.

On the blogA math trace that passes our checker is right about one time in threeWe checked a small model's math traces with a program. Accepted traces were right about one time in three; switching to the trace format cost accuracy.

More in Verifiable reasoning

Cite

@misc{bei2026measuring,
  title         = {Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Han, Yan and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability/}
}

Builds on

  1. Wei et al. (2022). Chain-of-thought prompting elicits reasoning in large language models.
  2. Lightman et al. (2024). Let’s verify step by step.
  3. Cobbe et al. (2021). Training verifiers to solve math word problems.
  4. Gao et al. (2023). PAL: Program-aided language models.
  5. Shao et al. (2024). DeepSeekMath: Pushing the limits of mathematical reasoning in open language models.