{
  "slug": "telegraph-reasoning",
  "title": "Telegraph Reasoning: Lintable Traces for Mechanically Verified Chain-of-Thought",
  "short": "TR",
  "line": "reasoning",
  "line_name": "Verifiable reasoning",
  "part": null,
  "status": "Preprint",
  "date": "2026-05-25",
  "authors": [
    "Sisong Bei",
    "Mikhail L Arbuzov",
    "Ziwei Dong",
    "Dmitri Kalaev",
    "Alexey Shvets"
  ],
  "abstract": "Reasoning models produce long chain-of-thought traces. When the final answer is wrong, the reader cannot easily point at the exact step that broke. Asking another large language model to find the error works in part, but the model disagrees with itself across reruns and misses certain error types entirely. We introduce Telegraph Reasoning, a discrete grammar for reasoning traces that lets a small rule-based linter and a symbolic algebra system check every step. The grammar has seven tags. The linter has eleven rules; four semantic rules use SymPy to check variable scope, units, numeric equations, and model-emitted verification claims. On a corpus of 266 traces with injected errors, the linter catches 99.5 percent of the errors. A self-verifying language model catches 66.3 percent. A frontier judge model catches 87.2 percent. The linter takes nine milliseconds per trace, returns the same answer on every run, and uses no language model at verification time.",
  "tldr": "If a model writes its math reasoning as tagged lines, a rule-based linter can recompute equations and check claims in SymPy instead of asking a model to reread them. It caught 195 of 196 injected errors; self-verification, about two thirds.",
  "pages": 9,
  "html": "https://telegrapher.ai/research/telegraph-reasoning/",
  "md": "https://telegrapher.ai/research/telegraph-reasoning.md",
  "reader": "https://telegrapher.ai/research/telegraph-reasoning/read/",
  "pdf": "https://telegrapher.ai/papers/telegraph-reasoning/telegraph-reasoning.pdf",
  "arxiv": null,
  "openreview": null,
  "post": "https://telegrapher.ai/blog/telegraph-reasoning/",
  "bibtex": "@misc{bei2026telegraph,\n  title         = {Telegraph Reasoning: Lintable Traces for Mechanically Verified Chain-of-Thought},\n  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},\n  year          = {2026},\n  note          = {Preprint},\n  url           = {https://telegrapher.ai/research/telegraph-reasoning/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "A seven-tag trace grammar lets a SymPy-backed linter catch 195 of 196 injected math errors"
    },
    {
      "label": "TL;DR",
      "text": "If a model writes its math reasoning as tagged lines, a rule-based linter can recompute equations and check claims in SymPy instead of asking a model to reread them. It caught 195 of 196 injected errors; self-verification, about two thirds."
    },
    {
      "label": "Method",
      "text": "Claude Sonnet 4.6 wrote Telegraph Reasoning traces for the full GSM8K (n = 1,319) and MATH500 (n = 500) test splits from five in-context examples, scored against three zero-shot baselines; separately, 266 traces (196 with one injected error from five categories, 70 clean) were checked by the structural rules alone, the full SymPy linter, Natural Program self-verification on deterministically translated traces, and Sonnet 4.6 as an LLM judge, the last two run three times at temperature 0.0 with a majority vote."
    },
    {
      "label": "Key result",
      "text": "Injected errors caught by the full linter: 99.5%; Injected errors caught by Natural Program self-verification: 66.3%; Median linter time per trace: 9 ms"
    },
    {
      "label": "Why it matters",
      "text": "Anyone who audits model reasoning with another model is paying for judgment on questions that are really computation."
    },
    {
      "label": "Limits",
      "text": "The adversarial corpus is synthetic."
    },
    {
      "label": "Status",
      "text": "Preprint, May 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/telegraph-reasoning/read/, PDF /papers/telegraph-reasoning/telegraph-reasoning.pdf"
    }
  ],
  "note": {
    "slug": "telegraph-reasoning",
    "claim_title": "A seven-tag trace grammar lets a SymPy-backed linter catch 195 of 196 injected math errors",
    "meta_description": "Telegraph Reasoning writes chain-of-thought as tagged lines a SymPy-backed linter can recompute. It caught 99.5% of injected errors; self-verification, 66.3%.",
    "tldr": "If a model writes its math reasoning as tagged lines, a rule-based linter can recompute equations and check claims in SymPy instead of asking a model to reread them. It caught 195 of 196 injected errors; self-verification, about two thirds.",
    "gist": "When a long chain of thought ends in a wrong answer, pointing at the step that broke is hard. A second language model asked to check the steps catches some errors. It also changes its verdict between reruns, and it struggles with a variable used before it is defined, a unit changed mid-derivation, or two equations that bind one name to different values. Telegraph Reasoning gives the trace a grammar of seven tags so that an eleven-rule linter can parse it; four of the rules hand the work to SymPy, checking variable scope, units, equation arithmetic and the model's own check claims. On 266 synthetic math traces, 196 of them carrying one injected error, the linter caught 195. Natural Program self-verification caught 66.3%, and Claude Sonnet 4.6 as a judge 87.2%. The linter took a median 9 ms per trace and gave the same verdict on every run, though it also flagged 7 of 70 clean traces. The errors were injected to match the rule categories, and on MATH500 the format trailed natural-language chain-of-thought by five to six points.",
    "method": "Claude Sonnet 4.6 wrote Telegraph Reasoning traces for the full GSM8K (n = 1,319) and MATH500 (n = 500) test splits from five in-context examples, scored against three zero-shot baselines; separately, 266 traces (196 with one injected error from five categories, 70 clean) were checked by the structural rules alone, the full SymPy linter, Natural Program self-verification on deterministically translated traces, and Sonnet 4.6 as an LLM judge, the last two run three times at temperature 0.0 with a majority vote.",
    "summary_html": [
      "A trace line reads <code>EQ[e1]: sold = total - eaten - baked</code>, and the next one asserts <code>CHECK[c1]: arith: e1.value == 9</code>. Nothing about that check needs a language model. A program looks up <code>total</code>, <code>eaten</code> and <code>baked</code> in the trace's <code>GIVEN</code> block, computes 16 − 3 − 4 itself, and compares the result with the model's 9; had the model asserted 8, it would report error <code>TE_E007</code> on that line. This is <em>Telegraph Reasoning</em>, a grammar of seven tags (<code>GIVEN</code>, <code>GOAL</code>, <code>STEP</code>, <code>EQ</code>, <code>CHECK</code>, <code>OPEN</code>/<code>RESOLVE</code>, <code>ANS</code>) with a linter of eleven rules. Seven rules check the trace's shape and need only the parse tree. The other four call SymPy, and each asks one question: is every symbol declared before use; does an equation's declared unit match the unit its right-hand side produces; where an equation's right-hand side resolves to a number, does it equal the declared left-hand side; and does every <code>CHECK</code> claim survive being recomputed from the trace? Claude Sonnet 4.6 wrote traces in this format for GSM8K and MATH500 from five worked examples. To test the checking, 196 traces each received one injected error (an arithmetic slip, a lost variable, an unsupported conclusion, a wrong unit or a contradicting equation) and were mixed with 70 clean ones; four checkers read all 266.",
      "The full linter caught 195 of the 196 errors. Natural Program self-verification, reading the same traces rendered into its prose format, managed 66.3%; Sonnet 4.6 prompted as a judge did better, at 87.2%. On their own the structural rules caught nothing, so every catch came from the four SymPy rules, and the gaps opened where checking is a computation. Wrong units are the starkest case: the linter caught 11 of 12, self-verification 2. Lost variables and contradicting equations went about half undetected by self-verification. Switching off the rule that re-evaluates equations cut detection from 99.5% to 78.1%, since 42 errors were caught by that rule and no other; it is the rule that catches a slip even when the model's own check repeats it. The linter's verdicts were identical on every run, at a median 9 ms per trace, while self-verification flipped on roughly 10% of traces across three runs at temperature 0.0. The linter's false-positive rate was also the highest of the four checkers, 10.0% of clean traces. As a format for writing, the tags held level with natural-language chain-of-thought on GSM8K, within a point, and trailed concise chain-of-thought by 5.80 points on MATH500."
    ],
    "key_numbers": [
      {
        "label": "Injected errors caught by the full linter",
        "value": "99.5%",
        "context": "195 of 196 single-error math traces; Wilson 95% CI 97.2% to 99.9%"
      },
      {
        "label": "Injected errors caught by Natural Program self-verification",
        "value": "66.3%",
        "context": "same 196 traces, Claude Sonnet 4.6, three runs with majority vote; Sonnet 4.6 as a judge caught 87.2%"
      },
      {
        "label": "Median linter time per trace",
        "value": "9 ms",
        "context": "95th percentile 81 ms; the 266-trace corpus lints in under 3 seconds on one CPU, with no model call"
      },
      {
        "label": "Linter false-positive rate",
        "value": "10.0%",
        "context": "7 of 70 clean traces flagged (7.1% under the paper's stricter audit reading); self-verification 2.9%, judge 5.7%",
        "bad": true
      },
      {
        "label": "MATH500 accuracy gap, tagged traces vs concise chain-of-thought",
        "value": "−5.80 pp",
        "context": "prompt-only Claude Sonnet 4.6, 95% CI −8.20 to −3.40; on GSM8K the format matched within a point",
        "bad": true
      }
    ],
    "editorial_html": [
      "Anyone who audits model reasoning with another model is paying for judgment on questions that are really computation. Does this equation evaluate to what the trace says? Was this symbol ever defined? Do the units on the two sides agree? Given a trace it can parse, a program answers these exactly. That is the trade Telegraph Reasoning offers: the model accepts a format, and the computational part of checking moves into a linter, where it behaves like a unit test. Same trace, same flags, same error codes and line numbers, run after run, which is something an audit pipeline can route on.",
      "Where the check claim goes is the real difference from self-verification. Natural Program asks the model whether its step holds; the linter recomputes it. The ablation complicates the design story, though. On this corpus the rule that re-evaluates equations did far more work than the rule that recomputes the model's check claims. It runs on every equation whose right-hand side resolves to a number, whether or not the model wrote a check line for it, and 42 errors were caught by it and no other; switching off the check rule cost under three points. The paper calls the equation rule its empirical workhorse and keeps the check rule as the conceptual line between the linter and self-verification. It presents the format as a substrate rather than a competitor, since reinforcement-learning methods that shorten traces leave them no more checkable, and the two could be combined.",
      "<em>Telegraph English</em> rewrites prose into atomic fact lines so that the compressed text doubles as an index. Both formats replace free prose with structured lines; one spends the structure on compression, the other on checking. A linter this fast and deterministic can also sit inside training. <em>Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning</em> needed a verifier quick enough to instrument every reward call, and used a rule-based trace linter with the detection rates reported here. <em>Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning</em> asks what this paper does not, whether a trace the checker accepts is also correct; for a small model trained by outcome-only reinforcement learning, an accepted trace was correct about one time in three."
    ],
    "limitations": "The adversarial corpus is synthetic. Each of the 196 errors was injected, one per trace, in five categories chosen to match the linter's rules, and the paper does not claim the 99.5% figure transfers to the errors models make on their own; what it expects to carry over is the mechanism, that re-evaluating an equation catches arithmetic mismatches the model's own self-verification cannot see. Everything is math, with logic, code and multi-trace settings untested, and one model, Claude Sonnet 4.6, wrote the traces, self-verified them and served as judge, so cross-model verification is open. Grammar v1 cannot yet express composite additive units, brute-force case enumeration or quantifier-rich logic, which is why one wrong-unit error got through. The accuracy cost is real: on MATH500 the format trailed concise chain-of-thought by 5.80 points, with five in-context examples against the baselines' none and no matched-shot comparison, and it spent more tokens per correct answer than concise chain-of-thought on both benchmarks. A preliminary fine-tune of a 2B model on 1,562 traces lost accuracy against the same model prompted for ordinary chain-of-thought. The linter's 10.0% false-positive rate on clean traces is the highest of the methods compared. And its rules test a trace's internal consistency (declarations, equations, units, check claims), not whether the trace matches the problem statement or the answer key.",
    "concept_terms": [
      "chain-of-thought verification",
      "reasoning trace linter",
      "SymPy symbolic verification",
      "Natural Program self-verification",
      "LLM-as-a-judge",
      "deterministic auditability"
    ],
    "references": [
      "Ling et al. (2023). Deductive verification of chain-of-thought reasoning.",
      "Hao et al. (2024). Training large language models to reason in a continuous latent space.",
      "Cheng and Van Durme (2024). Compressed chain of thought: Efficient reasoning through dense representations.",
      "Kontonis et al. (2026). MEMENTO: Teaching LLMs to manage their own context.",
      "Yang et al. (2026). Batched contextual reinforcement: A task-scaling law for efficient reasoning."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/telegraph-reasoning.txt",
    "verification": {
      "claims": [
        {
          "claim": "The grammar has seven tags: GIVEN, GOAL, STEP, EQ, CHECK, OPEN/RESOLVE, ANS",
          "verdict": "CONFIRMED",
          "source_quote": "The grammar has seven"
        },
        {
          "claim": "The tag list ends with OPEN/RESOLVE case-split markers and ANS",
          "verdict": "CONFIRMED",
          "source_quote": "(case-split markers), and ANS (the final answer)."
        },
        {
          "claim": "The linter has eleven rules; four use SymPy to check variable scope, units, equation arithmetic and model-emitted check claims",
          "verdict": "CONFIRMED",
          "source_quote": "The linter has eleven rules; four semantic rules use SymPy to check variable scope"
        },
        {
          "claim": "Seven structural rules check trace shape and need only the parse tree",
          "verdict": "CONFIRMED",
          "source_quote": "The seven structural rules check trace shape and"
        },
        {
          "claim": "Scope rule: every symbol declared in GIVEN or as the LHS of an earlier EQ",
          "verdict": "CONFIRMED",
          "source_quote": "be declared in given or be the LHS of an"
        },
        {
          "claim": "Unit rule: an equation's declared unit must match the unit of its right-hand side",
          "verdict": "CONFIRMED",
          "source_quote": "declared unit must match the unit of its"
        },
        {
          "claim": "Numeric rule: where an EQ right-hand side resolves to a number, it must equal the declared left-hand side",
          "verdict": "CONFIRMED",
          "source_quote": "RHS resolves to a numeric"
        },
        {
          "claim": "Check rule: the linter recomputes every CHECK claim from the trace",
          "verdict": "CONFIRMED",
          "source_quote": "check[id] the linter recomputes the"
        },
        {
          "claim": "Worked example lines EQ[e1]: sold = total - eaten - baked and CHECK[c1]: arith: e1.value == 9",
          "verdict": "CONFIRMED",
          "source_quote": "CHECK[c1]: arith: e1.value == 9"
        },
        {
          "claim": "Linter computes 16 - 3 - 4 = 9; a check asserting 8 would raise TE_E007 on that line",
          "verdict": "CONFIRMED",
          "source_quote": "would compute 9, disagree, and emit TE_E007 at"
        },
        {
          "claim": "Self-verification with a second LLM works in part but disagrees with itself across reruns",
          "verdict": "CONFIRMED",
          "source_quote": "but the model disagrees with itself across reruns"
        },
        {
          "claim": "Self-verification struggles with variables used before definition, changed units, and two equations binding one name to different values",
          "verdict": "CONFIRMED",
          "source_quote": "math: a variable used before it is defined, a unit"
        },
        {
          "claim": "Claude Sonnet 4.6 wrote traces on full GSM8K (n = 1,319) and MATH500 (n = 500) test splits",
          "verdict": "CONFIRMED",
          "source_quote": "full GSM8K test split (n = 1,319) and the full"
        },
        {
          "claim": "Telegraph Reasoning prompted with five in-context examples; the three baselines with none",
          "verdict": "CONFIRMED",
          "source_quote": "R5 (TE scratchpad with 5 in-context examples),"
        },
        {
          "claim": "Corpus of 266 traces: 196 adversarial, 70 clean",
          "verdict": "CONFIRMED",
          "source_quote": "266-trace adversarial corpus (196 adversarial, 70 clean)"
        },
        {
          "claim": "Each adversarial trace carries exactly one injected error from five categories (arithmetic slip, lost variable, unsupported conclusion, wrong unit, contradicting equation)",
          "verdict": "CONFIRMED",
          "source_quote": "Each adversarial trace contains exactly one injected error from one of five categories"
        },
        {
          "claim": "Most adversarial traces are programmatic perturbations of correct traces",
          "verdict": "CONFIRMED",
          "source_quote": "166 programmatic perturbations of correct R5 traces"
        },
        {
          "claim": "Natural Program self-verification ran on deterministically translated traces; Sonnet 4.6 served as LLM judge",
          "verdict": "CONFIRMED",
          "source_quote": "(Natural Program self-verify on deterministicallytranslated traces), M4 (LLM judge: Sonnet 4.6"
        },
        {
          "claim": "Model-based checkers run three times at temperature 0.0 with majority vote",
          "verdict": "CONFIRMED",
          "source_quote": "run three times per trace at temperature 0.0 with"
        },
        {
          "claim": "Self-verification and the judge both use Claude Sonnet 4.6",
          "verdict": "CONFIRMED",
          "source_quote": "TE traces and all LLM-judge baselines (M3,"
        },
        {
          "claim": "Full linter caught 195 of 196 injected errors",
          "verdict": "CONFIRMED",
          "source_quote": "195 of 196 injected errors (Wilson 95% CI"
        },
        {
          "claim": "Full linter TPR 99.5%",
          "verdict": "CONFIRMED",
          "source_quote": "SymPy verifier catches 99.5% of injected"
        },
        {
          "claim": "Wilson 95% CI 97.2% to 99.9%; Natural Program self-verification caught 66.3%",
          "verdict": "CONFIRMED",
          "source_quote": "[97.2%, 99.9%]). M3 catches 66.3%; M4 catches"
        },
        {
          "claim": "Sonnet 4.6 as a judge caught 87.2%",
          "verdict": "CONFIRMED",
          "source_quote": "87.2%. The M2-vs-M3 paired-bootstrap difference is +33.09 pp"
        },
        {
          "claim": "Structural rules alone caught nothing (M1 TPR 0.0%), so every catch came from the four SymPy rules (derived from Table 2)",
          "verdict": "CONFIRMED",
          "source_quote": "M1 (struct. only)"
        },
        {
          "claim": "The gaps open where checking is a computation",
          "verdict": "CONFIRMED",
          "source_quote": "categories where SymPy does work that naturallanguage verification cannot do reliably"
        },
        {
          "claim": "Wrong unit: linter 11 of 12, self-verification 2 of 12",
          "verdict": "CONFIRMED",
          "source_quote": "M2 catches 11 of 12 cases and M3 catches 2 of"
        },
        {
          "claim": "Self-verification caught 52.2% of lost variables (about half undetected)",
          "verdict": "CONFIRMED",
          "source_quote": "variable (52.2%), wrong unit (16.7%), and con-"
        },
        {
          "claim": "Self-verification caught 50.0% of contradicting equations (about half undetected)",
          "verdict": "CONFIRMED",
          "source_quote": "tradicting equation (50.0%)"
        },
        {
          "claim": "Disabling the equation re-evaluation rule cuts TPR from 99.5% to 78.1%",
          "verdict": "CONFIRMED",
          "source_quote": "from 99.5% to 78.1%)"
        },
        {
          "claim": "42 of 196 errors were caught by the equation rule alone, whether or not a check line exists",
          "verdict": "CONFIRMED",
          "source_quote": "rule 6 carries the largest empirical load: 42"
        },
        {
          "claim": "The equation rule catches a slip even when the model's own check agrees with it",
          "verdict": "CONFIRMED",
          "source_quote": "rule 6 catches a slip"
        },
        {
          "claim": "Disabling the check rule cost under three points (2.6 pp)",
          "verdict": "CONFIRMED",
          "source_quote": "drops by 2.6 pp"
        },
        {
          "claim": "The paper calls the equation rule the empirical workhorse and the check rule the conceptual distinction",
          "verdict": "CONFIRMED",
          "source_quote": "Rule 6 is the empirical workhorse on the corpus we evaluate"
        },
        {
          "claim": "Natural Program asks the model to verify a step; the linter recomputes it",
          "verdict": "CONFIRMED",
          "source_quote": "rather than asking an LLM to self-verify each"
        },
        {
          "claim": "Linter gives the same flags, error codes and line numbers on every run",
          "verdict": "CONFIRMED",
          "source_quote": "numbers on every run. M3 and M4 do not have"
        },
        {
          "claim": "Self-verification flipped its verdict on roughly 10% of traces across three runs",
          "verdict": "CONFIRMED",
          "source_quote": "flipped on roughly 10% of"
        },
        {
          "claim": "Determinism matters for audit pipelines that route flagged traces",
          "verdict": "CONFIRMED",
          "source_quote": "For audit pipelines that route flagged traces"
        },
        {
          "claim": "Median 9 ms per trace; 95th percentile 81 ms",
          "verdict": "CONFIRMED",
          "source_quote": "Latency. Median 9 ms per trace, 95th percentile"
        },
        {
          "claim": "The 266-trace corpus lints in under 3 seconds on one CPU",
          "verdict": "CONFIRMED",
          "source_quote": "266-trace corpus end-to-end takes under 3 seconds"
        },
        {
          "claim": "No language model at verification time",
          "verdict": "CONFIRMED",
          "source_quote": "answer on every run, and uses no language"
        },
        {
          "claim": "Linter raw FPR 10.0%, 7 of 70 clean traces flagged",
          "verdict": "CONFIRMED",
          "source_quote": "raw FPR is 10.0% (7 of 70 clean controls"
        },
        {
          "claim": "Stricter audit reading gives 7.1% FPR",
          "verdict": "CONFIRMED",
          "source_quote": "stricter audit-interpretation FPR of 7.1% drops"
        },
        {
          "claim": "Self-verification FPR 2.9% (Table 2 cell)",
          "verdict": "CONFIRMED",
          "source_quote": "2.9%"
        },
        {
          "claim": "Judge FPR 5.7% (Table 2 cell); so the linter's 10.0% is the highest of the four checkers",
          "verdict": "CONFIRMED",
          "source_quote": "5.7%"
        },
        {
          "claim": "MATH500 gap vs concise CoT -5.80 pp, 95% CI -8.20 to -3.40",
          "verdict": "CONFIRMED",
          "source_quote": "−5.80 pp (95% CI [−8.20, −3.40])"
        },
        {
          "claim": "On GSM8K the format matched NL chain-of-thought within a point (Table 1: 96.4% vs 97.0% / 96.8%)",
          "verdict": "CONFIRMED",
          "source_quote": "matches strong NL baselines on GSM8K and trails by"
        },
        {
          "claim": "On MATH500 the format trailed NL baselines by five to six points",
          "verdict": "CONFIRMED",
          "source_quote": "5–6 pp on MATH500"
        },
        {
          "claim": "Format spent more tokens per correct answer than concise CoT on both benchmarks (166/342 vs 55/148)",
          "verdict": "CONFIRMED",
          "source_quote": "55 / 148 tokens-per-correct (GSM8K / MATH500)"
        },
        {
          "claim": "Paper does not claim 99.5% transfers to naturally occurring errors",
          "verdict": "CONFIRMED",
          "source_quote": "We do not claim that the 99.5% TPR generalizes to naturally-occurring LLM failures"
        },
        {
          "claim": "Paper expects the mechanism to generalize: equation re-evaluation catches mismatches self-verification cannot see",
          "verdict": "CONFIRMED",
          "source_quote": "mechanism story is what we expect to generalize: rule 6 catches arithmetic mismatches that the"
        },
        {
          "claim": "Categories were chosen to match the linter's rule taxonomy",
          "verdict": "CONFIRMED",
          "source_quote": "categories matched to TE"
        },
        {
          "claim": "Math only; logic, code and multi-trace settings untested",
          "verdict": "CONFIRMED",
          "source_quote": "logic, code, and multi-trace settings is not"
        },
        {
          "claim": "One model generated and judged; cross-model verification is future work",
          "verdict": "CONFIRMED",
          "source_quote": "M4) use Claude Sonnet 4.6. Cross-model"
        },
        {
          "claim": "Grammar v1 gaps (composite additive units, brute-force enumeration, quantifier-rich logic) cause the one missed wrong-unit case",
          "verdict": "CONFIRMED",
          "source_quote": "quantifier-rich logic require v2 grammar extensions. M2 misses 1 of 12 wrong-unit cases"
        },
        {
          "claim": "Five-shot format against zero-shot baselines; matched-shot baselines are follow-up work",
          "verdict": "CONFIRMED",
          "source_quote": "TE uses 5-shot incontext examples; baselines use zero shots."
        },
        {
          "claim": "Preliminary 2B fine-tune on 1,562 traces lost accuracy against the same model prompted for standard CoT",
          "verdict": "CONFIRMED",
          "source_quote": "model (R5 sft : Qwen3.5-2B + TE-SFT on 1,562"
        },
        {
          "claim": "The rules check the trace's internal consistency, not the problem statement or answer key (derived from the rule definitions in 3.1 and Appendix C)",
          "verdict": "CONFIRMED",
          "source_quote": "every free symbol in eq, check, or ans must"
        },
        {
          "claim": "The format is a substrate, not a competitor, to RL methods that shorten traces",
          "verdict": "CONFIRMED",
          "source_quote": "The contribution is a substrate, not a competitor."
        },
        {
          "claim": "RL length-compression methods do not make the reasoning state verifiable",
          "verdict": "CONFIRMED",
          "source_quote": "None of these methods makes the reasoning state itself verifiable."
        },
        {
          "claim": "Reference: Ling et al. (2023) Deductive verification of chain-of-thought reasoning",
          "verdict": "CONFIRMED",
          "source_quote": "Deductive verification of chain-of-thought reasoning."
        },
        {
          "claim": "Reference: Hao et al. (2024) Training large language models to reason in a continuous latent space",
          "verdict": "CONFIRMED",
          "source_quote": "Training large language models to reason in a continuous latent space."
        },
        {
          "claim": "Reference: Cheng and Van Durme (2024) Compressed chain of thought",
          "verdict": "CONFIRMED",
          "source_quote": "Jeffrey Cheng and Benjamin Van Durme. 2024. Compressed chain of thought"
        },
        {
          "claim": "Reference: Kontonis et al. (2026) MEMENTO",
          "verdict": "CONFIRMED",
          "source_quote": "MEMENTO: Teaching"
        },
        {
          "claim": "Reference: Yang et al. (2026) Batched contextual reinforcement",
          "verdict": "CONFIRMED",
          "source_quote": "Batched contextual reinforcement: A taskscaling law for efficient reasoning."
        },
        {
          "claim": "[abstracts.yaml: telegraph-english] Telegraph English rewrites prose into atomic fact lines and the compressed output is also a semantic index",
          "verdict": "CONFIRMED",
          "source_quote": "it decomposes the input into atomic fact lines"
        },
        {
          "claim": "[abstracts.yaml: additive-process-rewards] That paper needed a verifier fast and deterministic enough to instrument every reward call and used a rule-based trace linter with 99.5 / 66.3 / 87.2 detection",
          "verdict": "CONFIRMED",
          "source_quote": "a verifier fast and deterministic enough to instrument every reward call"
        },
        {
          "claim": "[abstracts.yaml: answer-accuracy-and-trace-verifiability] For a small model trained by outcome-only RL, an accepted trace was correct about one time in three",
          "verdict": "CONFIRMED",
          "source_quote": "a trace the checker accepts is correct only about one time in"
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}