---
type: paper
slug: answer-accuracy-and-trace-verifiability
title: Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Yan Han
- Dmitri Kalaev
- Yanxin Zhang
- Alexey Shvets
date: '2026-08-25'
status: ICLR 2027 submission
line: Verifiable reasoning
pages: 26
html: https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability/
pdf: https://telegrapher.ai/papers/answer-accuracy-and-trace-verifiability/answer-accuracy-and-trace-verifiability.pdf
reader: https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability/read/
json: https://telegrapher.ai/api/papers/answer-accuracy-and-trace-verifiability.json
openreview: https://openreview.net/forum?id=xZMZw6Twv9
---

# Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

## Paper gist

- **Claim:** Math traces that passed a deterministic checker had the right answer about one time in three
- **TL;DR:** On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.
- **Method:** Outputs of Qwen3.5-4B checkpoints were scored twice, by a deterministic trace checker and by a separate reference-answer grader, in a panel of 25 trained checkpoints on a competition math suite at 2,048 tokens and in six seed pairs that trained the trace interface and free-form chain-of-thought with the same outcome-only reward, evaluated on 200 held-out GSM8K and MATH problems at a 1,024-token cap.
- **Key result:** Accuracy among checker-accepted outputs: 35.0%; Share of correct answers whose trace was accepted: 21.8%; Held-out accuracy change from adopting the trace interface: −15.6 pp
- **Why it matters:** Whoever filters or reports model outputs by a checker's verdict should read the joint table, not the verified fraction.
- **Limits:** The trained-model evidence covers one model, Qwen3.5-4B, on math, and acceptance means passing one checker's supported grammar and arithmetic rules; a checked claim is a specified test on a declared statement, not a check that the answer follows from the givens.
- **Status:** ICLR 2027 submission, August 2026
- **Read:** reader /research/answer-accuracy-and-trace-verifiability/read/, PDF /papers/answer-accuracy-and-trace-verifiability/answer-accuracy-and-trace-verifiability.pdf

## Abstract

Reasoning traces expose intermediate statements, but accepting a trace and obtaining a correct answer are distinct evaluation outcomes. We built a deterministic checker for a small model’s math reasoning traces and found that checkability and correctness come apart: a trace the checker accepts is correct only about one time in three. The model was trained by outcome-only reinforcement learning to write a compact trace language whose equations and checks a program evaluates. We measured answer accuracy and checker acceptance over the same outputs, using a checkpoint panel and seed-paired runs against free-form reasoning at matched token budgets. At matched training-example exposure, the trace interface raises acceptance by about 34 points and lowers accuracy by about 16 points relative to free-form reasoning. The directions of these differences are already present at the first measured checkpoint. Across the model’s own checkpoints, the accuracy–checkability trade-off we pre-registered did not confirm. Checkability and correctness are different quantities and should be measured separately.

## What we did and found

One trace spends 29 prose steps hunting for the value of a sum, tries 80? 81? and settles on 80. Then it writes EQ[eq_a]: S = 80 and CHECK[c1]: arith: eq_a.value == 80. Both claims check, the revised checker returns VERIFIED, and the grader marks the answer wrong, because the check confirms agreement with a value the trace asserted, not how that value follows from the givens. The trace language makes part of a model's output executable: GIVEN binds variables, GOAL names the target, STEP carries prose, EQ and CHECK expose claims a program evaluates, and ANS gives the answer. A deterministic checker's accepting verdict is called VERIFIED, and its share of all outputs the verified fraction. A separate grader takes another route through the same output: it extracts the final answer and compares it with the reference. Both were run over Qwen3.5-4B outputs in two studies. A panel of 25 trained checkpoints at 2,048 tokens, scored with the original v1 checker, supplies the joint distribution. Six seed pairs train the trace interface and free-form chain-of-thought from the same base model, data order and outcome-only reward, with no reward for acceptance.

On the panel's 17,025 pooled outputs, acceptance picked out a more accurate subset: 35.0% of accepted outputs were correct, against 18.38% of all outputs. A real filter, and a leaky one. Wrong-and-accepted outputs outnumbered correct-and-accepted ones, and the rejected column held 2,446 correct answers, so the filter kept 21.8% of the correct answers available. Adopting the trace interface changed the pool itself. Matched on training examples, it raised verified fraction by 34.1 points and lowered held-out accuracy by 15.6; matched on training tokens, the gaps were 35.1 and 19.1, and all six pairs shared both signs under both matches. Since free-form output does not target the trace language, its acceptance is zero and the acceptance gain equals the trace arm's own rate. Both gaps were already present at the first measured checkpoint. Within every trace run accuracy ended higher than it started, while acceptance moved up, down or not at all. Across checkpoints, the registered test for a negative accuracy–acceptance association did not confirm: the rank correlation was −0.087 against a target of −0.40 or lower, and none of 12 checkpoint pairs matched on accuracy and output length differed by the required five points of verified fraction.

## Key numbers

| Measure | Value |
|---|---|
| Accuracy among checker-accepted outputs | 35.0% |
| Share of correct answers whose trace was accepted | 21.8% |
| Held-out accuracy change from adopting the trace interface | −15.6 pp |
| Verified-fraction change from adopting the trace interface | +34.1 pp |
| Rank correlation of checkpoint accuracy with acceptance | −0.087 |

## Why it matters

Whoever filters or reports model outputs by a checker's verdict should read the joint table, not the verified fraction. A verified fraction says how many traces passed. It does not say how many answers were right, and on this panel the table puts two costs side by side: the error left among kept answers, and the correct answers thrown away. There is also a quieter problem. When the checker evaluates statements the model itself wrote, an accepting verdict can rest on self-assertion, as in the trace that bound its guessed answer to an equation and then checked the equation against the guess.

Choosing a reasoning format for its auditability is the second decision the results bear on. Under the same outcome-only reward, the format that made outputs checkable also produced fewer correct answers, and the accuracy gap persisted across the recorded checkpoints even as trace accuracy improved. The checkpoint result points the other way. Among models already writing traces, the data do not establish that more accurate checkpoints pass the checker less often, so the interface cost is not evidence of a general trade-off between accuracy and checkability.

Telegraph English rewrites a model's input context into compact symbolic statements; this interface structures the reasoning a model generates instead. Telegraph Reasoning measured how many injected errors a rule-based linter catches in reasoning traces, which is a different question from what a checker's verdict says about a model's own answers. And Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning finds trajectory-additive process rewards suppressed under group normalization; the paired runs here kept the reward on outcomes alone and measured acceptance beside accuracy.

## What this does not show

The trained-model evidence covers one model, Qwen3.5-4B, on math, and acceptance means passing one checker's supported grammar and arithmetic rules; a checked claim is a specified test on a declared statement, not a check that the answer follows from the givens. The interface comparison changes the trace language, prompting, decoding constraints and answer channel together, so it cannot say which component costs accuracy. Its first measurement comes at step 100, after training had begun. The study grew from three seed pairs to six after the first three results were visible, the paired means carry no population-level interval, and the trajectories are repeated readings of those six pairs rather than further replicates. The paired export does not record its checker version, and missing raw paired outputs and trainer state limit replay. The checkpoint test deviated from its registration too: it resampled checkpoints that share training lineages instead of problems, and matched pairs greedily rather than optimally.

## Blog post

[A math trace that passes our checker is right about one time in three](https://telegrapher.ai/blog/answer-accuracy-and-trace-verifiability.md)

## Cite

```bibtex
@misc{bei2026measuring,
  title         = {Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Han, Yan and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability/}
}
```
