telegrapher

A math trace that passes our checker is right about one time in three

Measuring Answer Accuracy and Trace Verifiability in Mathematical ReasoningThe paper: summary, reader and PDF

The trace below passed every check our program ran on it. The answer was wrong.

GOAL: S
...
STEP[s16]: Try integer values. 80? 81?
...
STEP[s25]: Actually, the sum is exactly 80.
...
EQ[eq_a]: S = 80
CHECK[c1]: arith: eq_a.value == 80
ANS: 80

One of our trained checkpoints wrote it for a MATH500 problem. Twenty-nine prose steps try identities, abandon them, estimate numerically and land on a round number. Then come two lines a program can evaluate: an equation binding the goal to 80, and a check that its value is 80. Both hold. The checker accepts the trace; the answer grader, comparing 80 with the reference, marks it wrong.

Nothing malfunctioned. The checker tested what the trace declared, and the trace declared that S was 80. The derivation that should connect the givens to that value stayed in prose, outside anything the checker counts.

The opposite failure exists too. Another checkpoint answered a different problem correctly, with 6, but its trace wrote RESOLVE[s4]: by eq_f without first opening an obligation called s4. That breaks a structural rule, and the checker rejects the trace.

One caution before any counting. Both verdicts come from v1.1, the revised version of our checker. The original v1 reached the opposite verdict on each: it rejected the S = 80 trace and accepted the other one. The panel frequencies further down were scored by v1.

Verdict and answer, then, can disagree in either direction. Two traces cannot say how often. If you use a checker’s verdict to filter answers, to choose a reasoning format or to compare trained models, frequency is the whole question, and the paper measures it.

Two readings of one output

The model writes in a compact trace language. Variables are bound with GIVEN, the target is named with GOAL, and the answer goes on an ANS line. Prose reasoning sits in STEP lines. The claims a program can test sit in EQ and CHECK lines, while OPEN and RESOLVE keep track of obligations. A deterministic checker evaluates those claims under the model’s own bindings; we call its accepting verdict VERIFIED, and the share of all outputs that receive it the verified fraction.

A separate grader takes the other route through the same output, extracting the final answer and comparing it with the reference. The problem itself reaches scoring through the grader, apart from anything the trace declares. Cross the two decisions and each output lands in one cell of a two-by-two table.

The two released versions of the checker differ in how they count. The original v1 rejects a trace when one of its rules detects a violation, and skips some expressions it cannot resolve. Its revision, v1.1, also takes an inventory of claims: VERIFIED needs at least one, all checked, while an empty inventory or an uncheckable claim yields a third verdict, partial. Re-scored on the same stored traces, the two versions let through about the same share, and v1.1 calls slightly more than half of it VERIFIED.

We measured answers and verdicts in two studies, each starting from Qwen3.5-4B. The first is a panel of 25 trained checkpoints, run on a competition math suite at 2,048 tokens and scored with v1. It supplies the joint distribution.

The second trains six seed pairs. In each pair, one run learns the trace interface and its partner learns ordinary free-form chain-of-thought. Both see the same data in the same order and are trained the same way: group-relative policy optimization, rewarded on the final answer alone. Passing the checker earns nothing. The pairs are evaluated on 200 held-out problems drawn from GSM8K and MATH, with generation capped at 1,024 tokens.

Acceptance is a real filter, and a leaky one

v1 verdict Correct answer Wrong answer
Accepted 683 1,271
Not accepted 2,446 12,625

Across 17,025 pooled checkpoint–problem outputs, 35.0% of accepted outputs had correct answers, against 18.38% of all outputs. Keeping the accepted column nearly doubles the hit rate, so the verdict is informative.

But it leaks, in both directions. Accepted-and-wrong outnumbers accepted-and-right nearly two to one, and the rejected column holds more than three times as many correct answers as the accepted one. Read from the answer’s side, a correct answer came with an accepted trace about one time in five.

Switching to traces costs accuracy from the first checkpoint

Filtering holds generation fixed; training the model to write traces changes it. Free-form chain-of-thought does not target the trace language, so its acceptance is zero at every measured checkpoint, and the acceptance gain from switching is simply the trace arm’s own rate. Accuracy is the column to watch.

Matched on Held-out accuracy Verified fraction
Training examples seen −15.6 pp +34.1 pp
Training tokens spent −19.1 pp +35.1 pp

Trace interface minus free-form, mean of six seed pairs.

All six pairs show lower accuracy and higher acceptance under both matches. Because the trace prompt is longer, each trace-arm step spends about 3.3 times as many training tokens, so matching on tokens sets the trace arm against free-form at a much earlier step. There the accuracy gap is wider.

The gaps did not open late. At the first measured checkpoint, step 100, every pair already had higher trace acceptance and lower trace accuracy, and both directions held through the last recorded checkpoint.

Inside the trace arm, accuracy finished above where it started in every run, while acceptance rose in some runs, fell in others and stayed flat in one. A trace run can improve over training and still trail its free-form partner.

The trade-off we registered was not confirmed

An earlier, hypothesis-forming analysis had found an association of −0.83 between accuracy and acceptance. The hypothesis we registered went further than the interface result: among checkpoints trained to write traces, the more accurate ones pass the checker less often. We tested it on fresh measurements, against two criteria set in advance.

Registered criterion Needed to confirm Observed
Rank correlation of checkpoint accuracy with acceptance −0.40 or lower, interval excluding zero −0.087 (95% interval −0.4576 to +0.3091)
Pairs matched on accuracy and output length that differ by five or more points of verified fraction, in the registered direction At least three None of 12, even ignoring direction; largest gap 3.98 points

The 25-checkpoint panel met neither criterion. The correlation landed close to zero, and its interval is wide enough to hold the threshold and positive values alike — the direction itself is unresolved. No matched pair reached the margin.

Two results that sound alike, then, are not. Switching formats cost accuracy. Among models already writing traces, the data do not establish that the more accurate ones are less checkable.

Read the joint table, not the pass rate

If you report a checker’s pass rate or filter by its verdict, put the joint table beside it. A verified fraction counts passing traces, not right answers, and here a format raised one while lowering the other under the same outcome reward. For a filter, the table names two costs: the errors left among the answers kept, and the correct answers thrown away.

Then there is what an accepting verdict establishes. A check that compares a value with the model’s own assertion of it will pass. In the S = 80 trace, two counted claims about one asserted number were enough for v1.1 to return VERIFIED where v1 had rejected the trace. Counting claims does make an empty or uncheckable inventory visible in the verdict. It does not yet tie the checked claims to what the problem gives and what it asks.

Telegraph English rewrites a model’s input context into compact symbolic statements; the interface here structures the reasoning a model generates instead. Telegraph Reasoning built a trace grammar with a rule-based linter and measured how many injected errors it catches, a different question from what a verdict says about a model’s own answers. And since the paired runs reward outcomes alone, a natural question is what rewarding checked steps would do; Why Additive Process Rewards Wash Out finds that the standard group-normalized pipeline suppresses trajectory-additive process rewards.

What is still open

The evidence comes from one model family at one scale, on math, and acceptance means passing one checker’s supported grammar and arithmetic rules. The interface comparison changes the trace language, prompt, decoding constraints and answer channel together; which of them costs accuracy needs component comparisons we have not run. The first measurement comes at step 100, after training began, so the trajectories cannot show how the gaps looked before training.

The paired study grew from three seed pairs to six after the first three results were visible. Its means carry no population-level interval, the export does not record which checker version scored it, and missing raw outputs and trainer state limit replay. The checkpoint test also deviated from its registration, resampling checkpoints that share training lineages instead of problems and matching pairs greedily rather than optimally. No gap reached the margin in either direction, so imposing the registered sign changes nothing there.

That leaves the question the S = 80 trace poses. A checker that followed the derivation from the givens to the answer would establish more than agreement with an assertion. Whether its verdict would then track the answer is a question for separate answer and acceptance measurements, taken side by side.

Read the paperMeasuring Answer Accuracy and Trace Verifiability in Mathematical ReasoningICLR 2027 submission, August 2026