{
  "slug": "what-survives-learned-symbolic-compression",
  "title": "What Survives Learned Symbolic Compression?",
  "short": "WSLSC",
  "line": "telegraph",
  "line_name": "Telegraph English",
  "part": null,
  "status": "ICLR 2027 submission",
  "date": "2026-08-25",
  "authors": [
    "Sisong Bei",
    "Mikhail L Arbuzov",
    "Ziwei Dong",
    "Dmitri Kalaev",
    "Yanxin Zhang",
    "Alexey Shvets"
  ],
  "abstract": "Lossy text compression for language-model pipelines is judged by reconstruction distance or by downstream task accuracy, and neither says whether the compressed representation still states what the source stated. We measure agreement with a constructed reference for grammar-constrained symbolic codes. The reference represents source facts as equations, and a fixed decoder recovers the equations a code encodes. An exact rule system compares their consequences independently of the code’s well-formedness checker. On controlled arithmetic micro-worlds, we evaluate learned translators from 70M to 2.8B parameters against structured controls and learned baselines at matched byte and token budgets. Passing the checker is not preserving the source: a self-consistent mistranslation passes it, and so do sampled learned-translation errors. Across byte budgets, the best learned trace translator trails the strongest compact structured control, with area under the closure-agreement curve of 0.67 against 0.86. Increasing translator size does not close this aggregate gap; the largest Pythia translator matches the best smaller translator’s saved byte-axis agreement scores. Source fidelity relative to the constructed reference, internal validity, and what a consumer model recovers are three different quantities and should be measured separately.",
  "tldr": "A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.",
  "pages": 22,
  "html": "https://telegrapher.ai/research/what-survives-learned-symbolic-compression/",
  "md": "https://telegrapher.ai/research/what-survives-learned-symbolic-compression.md",
  "reader": "https://telegrapher.ai/research/what-survives-learned-symbolic-compression/read/",
  "pdf": "https://telegrapher.ai/papers/what-survives-learned-symbolic-compression/what-survives-learned-symbolic-compression.pdf",
  "arxiv": null,
  "openreview": "https://openreview.net/forum?id=hEDqFXFy0v",
  "post": "https://telegrapher.ai/blog/what-survives-learned-symbolic-compression/",
  "bibtex": "@misc{bei2026what,\n  title         = {What Survives Learned Symbolic Compression?},\n  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},\n  year          = {2026},\n  note          = {ICLR 2027 submission},\n  url           = {https://telegrapher.ai/research/what-survives-learned-symbolic-compression/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "A learned symbolic code can pass its own checker and still change what the source said"
    },
    {
      "label": "TL;DR",
      "text": "A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open."
    },
    {
      "label": "Method",
      "text": "On constructed affine-arithmetic worlds (24,000 for training, 200 held out), trace translators trained with low-rank adapters on Pythia checkpoints from 70M to 2.8B parameters (three seeds each) and on two Qwen3 models were compared with Pythia 1.4B encoders trained to write canonical JSON, CCL-Core, CCL-Min and fixed-template prose, and with an irredundant-core oracle, at four byte and four token budgets, by closure agreement with the constructed reference, checker acceptance, and exact value recovery by Qwen3-Next-80B and GPT-OSS-20B."
    },
    {
      "label": "Key result",
      "text": "Byte-axis closure-AUC, best trace translator: 0.672; Closure agreement of a checker-valid mistranslation: 0.167; Checker-valid outputs among 40 sampled failures: 8–16"
    },
    {
      "label": "Why it matters",
      "text": "Anyone who lets a compressed code stand in for its source has a cheap test to hand: does the code parse and pass its checks?"
    },
    {
      "label": "Limits",
      "text": "The evidence comes from controlled affine arithmetic, whose reference language has no negation, modality, quantifiers, time or uncertainty."
    },
    {
      "label": "Status",
      "text": "ICLR 2027 submission, August 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/what-survives-learned-symbolic-compression/read/, PDF /papers/what-survives-learned-symbolic-compression/what-survives-learned-symbolic-compression.pdf"
    }
  ],
  "note": {
    "slug": "what-survives-learned-symbolic-compression",
    "claim_title": "A learned symbolic code can pass its own checker and still change what the source said",
    "meta_description": "On exact arithmetic, learned symbolic codes passed their own checker while changing source facts, and kept less per byte than a compact structured format.",
    "tldr": "A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.",
    "gist": "Compressed representations in language-model pipelines are usually judged by how well they reconstruct the source or by downstream task accuracy. Neither says whether the code still states what the source stated, and a code's own checker cannot say it either: a code that turns y = x + 3 into y = x + 4 consistently computes z = 12 and passes, though the source entails z = 10. Here the facts of controlled affine-arithmetic worlds are written as equations, a fixed decoder reads each code into the same equations, and the score is the Jaccard overlap of their exact consequences (closure agreement) at byte and token budgets. Sampled outputs of learned Pythia and Qwen3 translators also passed the checker while disagreeing with the source. Across byte budgets the best trace translator reached a closure-AUC of 0.672; CCL-Min, a compact format written by a 1.4B encoder, reached 0.860, and translators from 1B to 2.8B sat on a plateau. Counted in tokens, canonical JSON came out ahead instead. The domain is exact arithmetic; open text will need its own reference.",
    "method": "On constructed affine-arithmetic worlds (24,000 for training, 200 held out), trace translators trained with low-rank adapters on Pythia checkpoints from 70M to 2.8B parameters (three seeds each) and on two Qwen3 models were compared with Pythia 1.4B encoders trained to write canonical JSON, CCL-Core, CCL-Min and fixed-template prose, and with an irredundant-core oracle, at four byte and four token budgets, by closure agreement with the constructed reference, checker acceptance, and exact value recovery by Qwen3-Next-80B and GPT-OSS-20B.",
    "summary_html": [
      "Suppose a source says that x is 2, y is x plus 3, and z is 2 times y. A code that writes <code>y = x + 4</code> can compute <code>z = 12</code>, check that answer against its own equations, and pass; the source entails <code>z = 10</code>. Nothing inside the code can catch this. For controlled affine-arithmetic worlds, the comparison with the source can be made exact. The source facts become canonical equations, fixed before any code exists. A fixed decoder reads each code format into the same equation language, and exact rational rules expand both sides into a finite set of consequences: the explicit equations, the uniquely entailed values, and the uniquely entailed pairwise differences. The Jaccard overlap of the two sets is <em>closure agreement</em>. In the example the reference has seven consequences, the mistranslation shares two of the twelve in the union, and agreement drops to 0.167. Each code is scored at output allowances of 0.30, 0.45, 0.60 and 0.80 of the nonredundant source, counted once in bytes and once in tokens; an over-budget code scores zero, and the normalised area under the curve is <em>closure-AUC</em>.",
      "Learned translators produce checker-valid source errors as well. Among 40 sampled failing outputs from each of six translators (Pythia 410M to 2.8B and both Qwen3 models), 8 to 16 passed the checker. For Pythia 1B, Pythia 2.8B and Qwen3 0.6B none of the sampled failures was over budget, so the outputs the checker accepted there disagree with the source. Across byte budgets the best trace translator reached a closure-AUC of 0.672. CCL-Min, a compact structured format written by a Pythia 1.4B encoder, reached 0.860: the trace code trailed it under the tighter allowances and caught up near the largest. Size helped, then stopped helping. The 70M and 160M translators stayed near zero, 410M reached 0.654, and 1B, 1.4B and 2.8B sat between 0.668 and 0.672. Counting tokens instead of bytes reversed the order of canonical JSON, CCL-Core and CCL-Min, with canonical JSON ahead at 0.514 and CCL-Min behind at 0.426. Recovery by a consumer model is a separate measurement, taken on a 60-world subset: Qwen3-Next-80B recovered more values from canonical JSON and from the 1B trace code than from CCL-Min, and more from the oracle's explicit equations than from the source text."
    ],
    "key_numbers": [
      {
        "label": "Byte-axis closure-AUC, best trace translator",
        "value": "0.672",
        "context": "Pythia 1B and 2.8B; CCL-Min, written by a Pythia 1.4B encoder, reaches 0.860 over the same four byte budgets"
      },
      {
        "label": "Closure agreement of a checker-valid mistranslation",
        "value": "0.167",
        "context": "constructed fixture with y = x + 4 in place of y = x + 3; the exact code scores 1.000",
        "bad": true
      },
      {
        "label": "Checker-valid outputs among 40 sampled failures",
        "value": "8–16",
        "context": "per translator, for Pythia 410M to 2.8B and both Qwen3 models; samples are selected on failure, so they show the case exists, not its rate",
        "bad": true
      },
      {
        "label": "Token-axis closure-AUC, canonical JSON",
        "value": "0.514",
        "context": "ahead of the other structured controls on tokens; CCL-Min falls to 0.426 and trace translators reach at most 0.369"
      },
      {
        "label": "Exact value recovery by Qwen3-Next-80B, canonical JSON",
        "value": "0.703",
        "context": "against 0.471 for CCL-Min and 0.869 for the source text, on the 60-world consumer subset"
      }
    ],
    "editorial_html": [
      "Anyone who lets a compressed code stand in for its source has a cheap test to hand: does the code parse and pass its checks? The fixture and the sampled learned errors show what that test misses. A checker compares the code with itself, and a consistent mistranslation is, after all, consistent. Evidence that the content survived needs a reference on the source side, fixed before any code is written, plus a decoder that puts the code into the same terms.",
      "Which format wins depends on what is counted. CCL-Min comes out ahead per byte and canonical JSON per token, and the oracle, holding the correct facts by construction, still scores zero at the tightest byte allowance because its complete code does not fit. Past 1B, more translator capacity bought nothing on the byte axis, while a different target format did: a 1.4B encoder writing CCL-Min retained more per byte than any trace translator, the 2.8B included. An encoding comparison tells a pipeline something when its budget unit is the resource that pipeline actually spends.",
      "The trace format reuses the Telegraph English grammar and checker, so the question lands on that line of work. <em>Telegraph English</em> and <em>Context Compression Is Not One Thing</em> judged symbolic re-expression by whether a model could still answer questions over it; this paper adds the source-side check that question answering leaves out. Its instrument has a relative in <em>RippleKB</em>, which scores systems against a constructed dependency closure of the source units one edited fact affects."
    ],
    "limitations": "The evidence comes from controlled affine arithmetic, whose reference language has no negation, modality, quantifiers, time or uncertainty. Moving to open text needs a source representation and a consequence family chosen for that domain, and the finite signature already builds such choices into the score: logically equivalent equation sets can score differently. The checker-failure samples are selected on failure, so they show that accepted source errors occur, not how often, and the error categories that need manual inspection, offset changes like the fixture's among them, were left unassessed. The size plateau holds for one training and selection procedure, and the byte-axis curves of the 1B and 2.8B translators coincide exactly while their token-axis and checker measurements differ. Consumer recovery was measured on a 60-world subset across both budget axes, so its ordering cannot be set against closure agreement on the same codes; GPT-OSS-20B's recovery rounded to zero for every input but one under the 256-token response cap and rose once its reasoning effort was lowered, which makes recovery a function of decoding settings too. A planned unconstrained-prose control was dropped before training because 282 of 24,000 teacher assignments passed, too few to supervise it. The release holds aggregate tables, configuration and checkpoint hashes and the analysis code, but not raw outputs, weights, or the imported construction and decoding code, so the adapted CCL serialization and raw-output equality cannot be checked independently.",
    "concept_terms": [
      "closure agreement",
      "closure-AUC",
      "source fidelity",
      "checker validity",
      "learned symbolic compression",
      "consumer recovery"
    ],
    "references": [
      "Xu (2026). Semantic rate-distortion theory: Deductive compression and closure fidelity.",
      "Trukhina and Vashkelis (2026). Compress the context, keep the commitments: A formal framework for verifiable LLM context compression.",
      "Trukhina and Vashkelis (2026). SemanticZip: A pilot framework for lossy text compression with LLMs as semantic decompressors.",
      "Arbuzov et al. (2026). Telegraph English: Semantic prompt compression via structured symbolic rewriting.",
      "Pan et al. (2024). LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/what-survives-learned-symbolic-compression.txt",
    "verification": {
      "claims": [
        {
          "claim": "claim_title / tldr / meta_description: learned translators produce codes that pass the checker yet change the source",
          "verdict": "CONFIRMED",
          "source_quote": "Learned translators also produce internally valid codes that change their source."
        },
        {
          "claim": "gist: compressed representations are judged by reconstruction or downstream task accuracy, neither of which says whether the code still states what the source stated",
          "verdict": "CONFIRMED",
          "source_quote": "is judged by reconstruction distance or by downstream task accuracy, and neither says whether the compressed representation still states what the source stated"
        },
        {
          "claim": "Worked example: source x = 2, y = x + 3, z = 2y; the code y = x + 4 computes and checks z = 12 while the source entails z = 10",
          "verdict": "CONFIRMED",
          "source_quote": "Its internal checks succeed even though the source entails z = 10."
        },
        {
          "claim": "The checker tests the code against itself only",
          "verdict": "CONFIRMED",
          "source_quote": "it tests the code’s grammar and supported internal constraints using only the code"
        },
        {
          "claim": "The reference (canonical equations) is fixed before any code exists",
          "verdict": "CONFIRMED",
          "source_quote": "The reference projection is determined before any compressed code exists."
        },
        {
          "claim": "Each code format has a fixed decoder into the same equation language",
          "verdict": "CONFIRMED",
          "source_quote": "Each supported code format has a fixed decoder into this equation language."
        },
        {
          "claim": "Consequence set = explicit equations, uniquely entailed values, uniquely entailed pairwise differences, by exact rational rules",
          "verdict": "CONFIRMED",
          "source_quote": "explicit canonical equations, uniquely entailed variable values, and uniquely entailed pairwise differences"
        },
        {
          "claim": "Jaccard overlap of the two consequence sets is closure agreement",
          "verdict": "CONFIRMED",
          "source_quote": "Their Jaccard overlap is closure agreement"
        },
        {
          "claim": "The example's reference has seven consequences",
          "verdict": "CONFIRMED",
          "source_quote": "The reference in Figure 1 has seven distinct consequences"
        },
        {
          "claim": "Mistranslation shares two of twelve in the union; agreement 0.167 (key number; recomputed 2/12 = 0.167)",
          "verdict": "CONFIRMED",
          "source_quote": "Source agreement: 2/12 = 0.167"
        },
        {
          "claim": "Exact code scores 1.000 (key number context)",
          "verdict": "CONFIRMED",
          "source_quote": "Exact-code fixture: checker accepts; closure agreement 1.000"
        },
        {
          "claim": "Allowances 0.30, 0.45, 0.60, 0.80 of the nonredundant source",
          "verdict": "CONFIRMED",
          "source_quote": "We evaluate nominal fractions 0.30, 0.45, 0.60, and 0.80 of each world’s nonredundant source rendering"
        },
        {
          "claim": "Counted once in bytes and once in tokens (four byte and four token budgets), separate outputs per axis",
          "verdict": "CONFIRMED",
          "source_quote": "Byte and token conditions generate separate outputs under their respective allowances."
        },
        {
          "claim": "An over-budget code scores zero",
          "verdict": "CONFIRMED",
          "source_quote": "An over-budget code stays in the evaluation denominator and receives zero matched-budget agreement."
        },
        {
          "claim": "closure-AUC is the normalised area under the agreement curve",
          "verdict": "CONFIRMED",
          "source_quote": "Every budget point contributes to this normalized area"
        },
        {
          "claim": "method: 24,000 training worlds, 200 held-out worlds",
          "verdict": "CONFIRMED",
          "source_quote": "Training uses 24,000 worlds; held-out reporting uses 200 worlds"
        },
        {
          "claim": "method: trace translators with low-rank adapters on Pythia 70M to 2.8B, three seeds",
          "verdict": "CONFIRMED",
          "source_quote": "Trace translators use low-rank adapters on Pythia checkpoints from 70M to 2.8B parameters, with three seeds labelled A, B, and C"
        },
        {
          "claim": "method: two Qwen3 translators (Qwen3 0.6B and 1.7B in Tables 5, 6, 8)",
          "verdict": "CONFIRMED",
          "source_quote": "Additional Qwen translators provide a second-family comparison"
        },
        {
          "claim": "method / summary: canonical JSON, CCL-Core, CCL-Min and fixed-template prose written by trained Pythia 1.4B encoders",
          "verdict": "CONFIRMED",
          "source_quote": "canonical JSON, CCL-Core, CCL-Min, and fixed-template prose—use trained Pythia 1.4B encoders"
        },
        {
          "claim": "method: irredundant-core oracle reads the construction directly",
          "verdict": "CONFIRMED",
          "source_quote": "The irredundant-core oracle reads the construction directly and serializes its correct core facts."
        },
        {
          "claim": "method / limitations: consumers Qwen3-Next-80B and GPT-OSS-20B with a 256-token response cap",
          "verdict": "CONFIRMED",
          "source_quote": "The consumers are Qwen3-Next-80B and GPT-OSS-20B, with a registered 256-token response cap"
        },
        {
          "claim": "40 sampled failing outputs per translator",
          "verdict": "CONFIRMED",
          "source_quote": "For each translator, the analysis selects 40 failing output records"
        },
        {
          "claim": "8 to 16 of the 40 passed the checker, for the six translators in Table 1 (Pythia 410M, 1B, 1.4B, 2.8B, Qwen3 0.6B, 1.7B) (key number 8–16, scope corrected)",
          "verdict": "CONFIRMED",
          "source_quote": "The samples contain 8–16 linter-valid outputs per model."
        },
        {
          "claim": "Pythia 1B, Pythia 2.8B and Qwen3 0.6B samples have no over-budget records, so their accepted failures disagree with the source",
          "verdict": "CONFIRMED",
          "source_quote": "The Pythia 1B, Pythia 2.8B, and Qwen 0.6B samples contain no over-budget records; their accepted failures therefore have imperfect source agreement."
        },
        {
          "claim": "Samples are selected on failure: they show the case exists, not its rate",
          "verdict": "CONFIRMED",
          "source_quote": "The diagnostic samples identify failure modes rather than estimate their prevalence across all outputs."
        },
        {
          "claim": "limitations (added): error categories needing manual inspection, coefficient or offset among them, were left unassessed",
          "verdict": "CONFIRMED",
          "source_quote": "Five categories require manual inspection and remain unassessed: entity substitution, sign or direction, coefficient or offset"
        },
        {
          "claim": "Best trace translator byte closure-AUC 0.672 vs CCL-Min 0.860 (key number)",
          "verdict": "CONFIRMED",
          "source_quote": "0.672 against 0.860 for CCL-Min"
        },
        {
          "claim": "Best trace value is shared by Pythia 1B and 2.8B (key number context)",
          "verdict": "CONFIRMED",
          "source_quote": "The 1B and 2.8B models both reach 0.672"
        },
        {
          "claim": "Trace code trails CCL-Min under the tighter allowances and approaches it at the largest",
          "verdict": "CONFIRMED",
          "source_quote": "under the tighter allowances and approaches it at the largest allowance"
        },
        {
          "claim": "70M and 160M near zero; 410M reaches 0.654",
          "verdict": "CONFIRMED",
          "source_quote": "The 70M and 160M systems remain near zero, while the 410M translator reaches 0.654."
        },
        {
          "claim": "1.4B reaches 0.668; 1B, 1.4B and 2.8B between 0.668 and 0.672",
          "verdict": "CONFIRMED",
          "source_quote": "the intervening 1.4B model reaches 0.668"
        },
        {
          "claim": "Translators from 1B to 2.8B sit on a plateau; larger translators leave the gap open",
          "verdict": "CONFIRMED",
          "source_quote": "The larger Pythia systems share a plateau in byte-axis closure agreement."
        },
        {
          "claim": "Token axis: canonical JSON 0.514 ahead, CCL-Min 0.426 (key number)",
          "verdict": "CONFIRMED",
          "source_quote": "Canonical JSON leads that axis at 0.514, compared with 0.426 for CCL-Min"
        },
        {
          "claim": "Token axis: trace translators reach at most 0.369 (key number context)",
          "verdict": "CONFIRMED",
          "source_quote": "trace translators reach at most 0.369"
        },
        {
          "claim": "Tokens reverse the order of JSON, CCL-Core, CCL-Min (Table 2: bytes 0.624 / 0.776 / 0.860, tokens 0.514 / 0.471 / 0.426; recomputed, exact reversal)",
          "verdict": "CONFIRMED",
          "source_quote": "Token accounting changes the ordering among structured controls."
        },
        {
          "claim": "editorial: CCL-Min leads per byte, JSON per token",
          "verdict": "CONFIRMED",
          "source_quote": "CCL-Min leads aggregate agreement per byte, while JSON leads under token accounting."
        },
        {
          "claim": "editorial: a 1.4B encoder writing CCL-Min (0.860) retained more per byte than every trace translator incl. 2.8B (max 0.672); past 1B, capacity bought nothing on bytes (1B 0.672, 1.4B 0.668, 2.8B 0.672, Table 5)",
          "verdict": "CONFIRMED",
          "source_quote": "Increasing translator size improves the weakest trace systems but does not close the aggregate gap"
        },
        {
          "claim": "editorial: oracle scores zero at the tightest byte allowance because its complete code does not fit",
          "verdict": "CONFIRMED",
          "source_quote": "Its byte-axis agreement is zero at 0.30 and perfect from 0.60, when its complete code fits."
        },
        {
          "claim": "editorial: an encoding comparison is useful when its budget unit is the resource the pipeline spends",
          "verdict": "CONFIRMED",
          "source_quote": "An encoding comparison becomes useful when its budget matches the resource the intended pipeline must conserve."
        },
        {
          "claim": "Consumer recovery measured on a 60-world subset (key number context)",
          "verdict": "CONFIRMED",
          "source_quote": "consumer evaluation uses a fixed subset of 60 worlds with 240 renderings"
        },
        {
          "claim": "Qwen3-Next-80B recovers more from canonical JSON (0.703) than from CCL-Min (0.471); source text 0.869 (Table 4) (key number)",
          "verdict": "CONFIRMED",
          "source_quote": "Qwen3-Next-80B recovers values more accurately from canonical JSON than from CCL-Min"
        },
        {
          "claim": "Qwen3-Next-80B recovers more from the 1B trace code (0.586) than from CCL-Min",
          "verdict": "CONFIRMED",
          "source_quote": "Canonical JSON and the displayed 1B trace code exceed CCL-Min’s recovery point estimate"
        },
        {
          "claim": "Qwen3-Next-80B recovers more from the oracle's equations (0.907) than from the source text (0.869)",
          "verdict": "CONFIRMED",
          "source_quote": "The consumer’s recovery point estimate is higher for the oracle’s explicit equations than for the rendered source."
        },
        {
          "claim": "limitations: consumer ordering cannot be set against closure agreement on the same codes",
          "verdict": "CONFIRMED",
          "source_quote": "relating its ordering to closure agreement would require both measurements on the same codes and worlds"
        },
        {
          "claim": "limitations: GPT-OSS-20B recovery rounds to zero for every input but one under the 256-token cap",
          "verdict": "CONFIRMED",
          "source_quote": "Primary GPT-OSS-20B recovery rounds to zero for all but the strong-encoder input"
        },
        {
          "claim": "limitations: lowering GPT-OSS-20B's reasoning effort recovered some values",
          "verdict": "CONFIRMED",
          "source_quote": "The secondary GPT-OSS condition changes reasoning effort to low"
        },
        {
          "claim": "limitations: reference language has no negation, modality, quantifiers, time or uncertainty",
          "verdict": "CONFIRMED",
          "source_quote": "The reference language excludes negation, modality, quantifiers, time, and uncertainty."
        },
        {
          "claim": "limitations: open text needs its own source representation and consequence family",
          "verdict": "CONFIRMED",
          "source_quote": "Extending it to open text requires a source representation and consequence family appropriate to that domain"
        },
        {
          "claim": "limitations: logically equivalent equation sets can score differently",
          "verdict": "CONFIRMED",
          "source_quote": "logically equivalent affine bases can have different scored sets"
        },
        {
          "claim": "limitations: plateau holds for one training and selection procedure",
          "verdict": "CONFIRMED",
          "source_quote": "The plateau is therefore a result about saved agreement scores under this training setup."
        },
        {
          "claim": "limitations: 1B and 2.8B byte-axis curves coincide while token-axis and checker results differ (Table 5 per-budget cells identical)",
          "verdict": "CONFIRMED",
          "source_quote": "The recorded byte-axis curves for 1B and 2.8B coincide, while token-axis and validity measurements differ."
        },
        {
          "claim": "limitations: unconstrained-prose control dropped before training; 282 of 24,000 teacher assignments passed",
          "verdict": "CONFIRMED",
          "source_quote": "Only 282 of 24,000 assignments passed, leaving too little supervision to train the matched unconstrained-prose control"
        },
        {
          "claim": "limitations: release omits raw outputs, weights, imported construction/decoding code; CCL serialization and raw-output equality not checkable",
          "verdict": "CONFIRMED",
          "source_quote": "The exact adapted CCL serialization is therefore unavailable for independent inspection."
        },
        {
          "claim": "editorial: the trace format reuses the Telegraph English grammar and checker",
          "verdict": "CONFIRMED",
          "source_quote": "The TE grammar, structural checker, full checker, and training-harness patterns are reused infrastructure."
        },
        {
          "claim": "editorial (other paper, abstracts.yaml): Telegraph English was judged by question answering",
          "verdict": "CONFIRMED",
          "source_quote": "We evaluate TE on 4,081 question-answer pairs from LongBench-v2"
        },
        {
          "claim": "editorial (other paper, abstracts.yaml): Context Compression Is Not One Thing judged symbolic re-expression by multi-hop QA",
          "verdict": "CONFIRMED",
          "source_quote": "We study context compression for multi-hop question answering with small language models."
        },
        {
          "claim": "editorial (other paper, abstracts.yaml): RippleKB scores against a constructed dependency closure of the source units one edited fact affects",
          "verdict": "CONFIRMED",
          "source_quote": "The proposed answer key is a hidden dependency closure"
        },
        {
          "claim": "reference present in bibliography: Xu (2026c)",
          "verdict": "CONFIRMED",
          "source_quote": "Semantic rate-distortion theory: Deductive compression and closure fidelity"
        },
        {
          "claim": "reference present in bibliography: Trukhina and Vashkelis (2026a)",
          "verdict": "CONFIRMED",
          "source_quote": "Compress the context, keep the commitments: A formal framework for verifiable LLM context compression"
        },
        {
          "claim": "reference present in bibliography: Trukhina and Vashkelis (2026b)",
          "verdict": "CONFIRMED",
          "source_quote": "SemanticZip: A pilot framework for lossy text compression with LLMs as semantic decompressors"
        },
        {
          "claim": "reference present in bibliography: Arbuzov et al. (2026)",
          "verdict": "CONFIRMED",
          "source_quote": "Telegraph English: Semantic prompt compression via structured symbolic rewriting"
        },
        {
          "claim": "reference present in bibliography: Pan et al. (2024)",
          "verdict": "CONFIRMED",
          "source_quote": "LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression"
        },
        {
          "claim": "CORRECTED (summary_html[1]): 'Learned translators make the fixture's mistake on their own' implied the learned errors are the fixture's offset change; Appendix F leaves the coefficient/offset category unassessed. Reworded to 'produce checker-valid source errors as well'.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (summary_html[1], key_numbers[2]): '8 to 16 passed the checker' per translator was unscoped; Table 1 covers six translators (Pythia 410M to 2.8B, two Qwen3), and Table 15 gives 5 and 4 linter-valid records for 160M and 70M. Scoped to the six.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (summary_html[1]): 'accepted outputs simply fall short of the source' narrowed imperfect source agreement to omission; reworded to 'disagree with the source'.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (summary_html[1]): 'turned the ranking of the structured codes upside down' holds only for JSON / CCL-Core / CCL-Min (structured prose, byte 0.749 and token 0.420, breaks a full reversal). Named the three codes.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (summary_html[1]): 'A consumer model produced yet another ordering' set the consumer ranking against the closure rankings, which the paper says needs both measures on the same codes and worlds. Reworded as a separate measurement.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}