{
  "slug": "telegraph-english",
  "title": "Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting",
  "short": "TE",
  "line": "telegraph",
  "line_name": "Telegraph English",
  "part": null,
  "status": "NeurIPS 2026 submission",
  "date": "2026-05-04",
  "authors": [
    "Mikhail L Arbuzov",
    "Sisong Bei",
    "Ziwei Dong",
    "Dmitri Kalaev",
    "Alexey Shvets",
    "Lee Mosbacker"
  ],
  "abstract": "We introduce Telegraph English (TE), a prompt-compression protocol that rewrites natural language into a symbol-rich, formally-structured dialect. Where tokendeletion methods such as LLMLingua-2 train a classifier to delete low-importance tokens at a fixed ratio, TE performs a full semantic rewrite: it decomposes the input into atomic fact lines, substitutes verbose phrases with ∼40 logical and relational symbols, and lets the compression ratio adapt to each document’s information density. A consequence of the line-structure rule is that compression and semantic chunking become the same operation—each output line is an independently addressable fact, so the compressed representation is simultaneously a semantic index. We evaluate TE on 4,081 question-answer pairs from LongBench-v2 across five OpenAI models and two difficulty levels. At roughly 50% token reduction, TE preserves 99.1% accuracy on key facts with GPT-4.1 and outperforms LLMLingua-2 at matched compression ratios on every model and task tested. The gap widens on smaller models—up to 11 percentage points on fine-detail tasks—suggesting that explicit relational structure compensates for limited model capacity. We release the grammar specification, compression prompt, benchmark data, and reference implementation.",
  "tldr": "Rewriting text so each line holds one claim, with explicit symbols for cause and contrast, cut tokens by about two fifths and lost fewer answers than LLMLingua-2's token deletion, by the widest margin on fine details.",
  "pages": 18,
  "html": "https://telegrapher.ai/research/telegraph-english/",
  "md": "https://telegrapher.ai/research/telegraph-english.md",
  "reader": "https://telegrapher.ai/research/telegraph-english/read/",
  "pdf": "https://telegrapher.ai/papers/telegraph-english/telegraph-english.pdf",
  "arxiv": "2605.04426",
  "openreview": "https://openreview.net/forum?id=VNQstvpecn",
  "post": "https://telegrapher.ai/blog/telegraph-english/",
  "bibtex": "@misc{arbuzov2026telegraph,\n  title         = {Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting},\n  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey and Mosbacker, Lee},\n  year          = {2026},\n  note          = {NeurIPS 2026 submission},\n  eprint        = {2605.04426},\n  archivePrefix = {arXiv},\n  url           = {https://telegrapher.ai/research/telegraph-english/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "Rewriting prompts one fact per line loses fewer fine-detail answers than deleting tokens"
    },
    {
      "label": "TL;DR",
      "text": "Rewriting text so each line holds one claim, with explicit symbols for cause and contrast, cut tokens by about two fifths and lost fewer answers than LLMLingua-2's token deletion, by the widest margin on fine details."
    },
    {
      "label": "Method",
      "text": "339 LongBench-v2 documents were split into 4,081 chunks, compressed with TE (o4-mini running the v5 grammar prompt) and with LLMLingua-2 at 50% retention (33% as a more aggressive setting), and tested with four-option questions written by GPT-4.1 (4,081 on key facts, 801 adversarial fine-detail items) and answered by GPT-4.1, GPT-4o, GPT-4o-mini and GPT-4.1-nano."
    },
    {
      "label": "Key result",
      "text": "Key-fact accuracy on TE text, GPT-4.1: 99.1%; Mean TE compression ratio: 0.585; Fine-detail accuracy drop on TE text, GPT-4o: −3.1 pp"
    },
    {
      "label": "Why it matters",
      "text": "Teams running retrieval and agent pipelines on smaller, cheaper reader models have the most at stake."
    },
    {
      "label": "Limits",
      "text": "Every result here is static: compress once, read once, score."
    },
    {
      "label": "Status",
      "text": "NeurIPS 2026 submission, May 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/telegraph-english/read/, PDF /papers/telegraph-english/telegraph-english.pdf, arXiv:2605.04426 https://arxiv.org/abs/2605.04426"
    }
  ],
  "note": {
    "slug": "telegraph-english",
    "claim_title": "Rewriting prompts one fact per line loses fewer fine-detail answers than deleting tokens",
    "meta_description": "Telegraph English compresses prompts by rewriting them as one-fact lines with explicit symbols. On LongBench-v2 QA it lost fewer answers than LLMLingua-2.",
    "tldr": "Rewriting text so each line holds one claim, with explicit symbols for cause and contrast, cut tokens by about two fifths and lost fewer answers than LLMLingua-2's token deletion, by the widest margin on fine details.",
    "gist": "Prompt compressors such as LLMLingua-2 save tokens by deleting the ones a classifier scores as unimportant. What comes out has no structure. Dropped connectives leave the reading model to guess how the surviving fragments relate, and a lone number can look dispensable right up until a question turns on it. Telegraph English (TE) rewrites instead. An LLM breaks the text into lines that each carry one claim, swaps verbose phrasing for roughly 40 logical and relational symbols, and lets the compression ratio follow the text's information density; because each line is a self-contained fact, the compressed text doubles as a chunked index. On 4,081 multiple-choice questions over LongBench-v2 chunks, at a mean compression ratio of 0.585, GPT-4.1 stayed within a point of its original accuracy, and TE scored at or above LLMLingua-2 at 50% retention for every reader model tested: thinly on headline facts, more clearly on 801 adversarial fine-detail questions. The context-management uses of the line structure are argued rather than benchmarked, and every reader model was a closed OpenAI model.",
    "method": "339 LongBench-v2 documents were split into 4,081 chunks, compressed with TE (o4-mini running the v5 grammar prompt) and with LLMLingua-2 at 50% retention (33% as a more aggressive setting), and tested with four-option questions written by GPT-4.1 (4,081 on key facts, 801 adversarial fine-detail items) and answered by GPT-4.1, GPT-4o, GPT-4o-mini and GPT-4.1-nano.",
    "summary_html": [
      "A sentence reporting that Johnson and colleagues' machine-learning diagnostics raised early detection by 27.5% and cut false positives by about 12% comes out as <code>ML → MEDICAL-DIAGNOSTICS: EARLY-DETECTION+27.5% ∧ FALSE-POSITIVE-12% [JOHNSON:2023]</code>. Sixty-eight tokens become fourteen, and the cause, both figures and the citation each keep a slot of their own. That rewrite is <em>Telegraph English</em> (TE). Its 430-line grammar doubles as the compressor's system prompt and sets out what it wants from the LLM: one claim per line; a fixed vocabulary of about 40 symbols in place of verbose phrasing (<code>→</code> for causation, <code>∴</code> for a conclusion, <code>VS</code> for contrast); tags for scope and modality; and fidelity ranked above brevity. To test it, 4,081 LongBench-v2 chunks were compressed with TE and with LLMLingua-2, and reader models answered the same four-option questions on the original and on each compressed version.",
      "TE's output averaged 0.585 of the original token count. Verbose narrative shrank hard, while a few short, dense chunks came out longer than they went in, because the grammar will not drop information to save tokens. On key facts the two methods were close: GPT-4.1 scored 99.1% on TE text, TE and LLMLingua-2 at 50% retention landed within a tenth of a point of each other on GPT-4.1 and GPT-4.1-nano, and TE edged ahead on GPT-4o-mini. Fine-detail questions cost both methods two and a half to three times as much accuracy on GPT-4o-mini, the one reader run on both suites, and there the margin opened. GPT-4o lost about half as many points reading TE as reading LLMLingua-2's output, and against LLMLingua-2 at 33% retention TE's lead reached roughly 11 points on GPT-4o-mini. TE has failures of its own, and they cluster on dates, units and qualifiers: in one legal chunk, <em>no later than 30 calendar days</em> became <code>30D</code>, and the question asked whether the days were calendar or business days."
    ],
    "key_numbers": [
      {
        "label": "Key-fact accuracy on TE text, GPT-4.1",
        "value": "99.1%",
        "context": "4,081 questions; 1.000 on the original text, 0.990 with LLMLingua-2 at 50% retention"
      },
      {
        "label": "Mean TE compression ratio",
        "value": "0.585",
        "context": "compressed over original tokens across 4,081 chunks, a 41.5% reduction; range 0.13 to 1.57"
      },
      {
        "label": "Fine-detail accuracy drop on TE text, GPT-4o",
        "value": "−3.1 pp",
        "context": "801 adversarial questions; LLMLingua-2 at 50% retention dropped −6.3 pp"
      },
      {
        "label": "TE lead over LLMLingua-2 at 33% retention, fine details, GPT-4o-mini",
        "value": "11 pp",
        "context": "roughly; LLMLingua-2 at 33% fell 21 pp from baseline; not a matched-ratio comparison"
      },
      {
        "label": "Key-fact items correct on the original but wrong on TE, GPT-4.1-nano",
        "value": "4.6%",
        "context": "187 of 4,081 items; failures cluster on dates, units, qualifiers and numerical relationships",
        "bad": true
      }
    ],
    "editorial_html": [
      "Teams running retrieval and agent pipelines on smaller, cheaper reader models have the most at stake. That is where compression pays for itself, and where deletion's losses ran largest in these tests. What changes is who does the reconstruction. A deletion classifier scores tokens for importance, and a number or a qualifier can look dispensable beside the prose around it, so the reader has to rebuild relationships from whatever fragments survive. TE does that rebuilding once, at compression time: causation, contrast and conclusion arrive as explicit symbols, and the grammar requires each number to stay attached to its unit.",
      "The bigger change is to what compressed text is. Since each TE line is one fact under a tagged heading, a context-assembly step can keep the relevant lines, collapse a section to its heading, or drop it, using string operations and no further model call. One expensive rewrite; after that, cheap edits for as long as the text stays in use. The paper offers this compress-once, manage-continuously principle as an architectural argument.",
      "Three of this paper's loose ends point to other work in the series. The LLMLingua-2 match is approximate, with TE's mean ratio of 0.585 above the 50% retention setting, and <em>Evaluating Relational Context Compression at Realized Token Budgets</em> makes realized budgets the basis of comparison. Multiple-choice accuracy shows whether a reader can still pick the right answer, which tells only indirectly whether the compressed text states what its source stated; <em>What Survives Learned Symbolic Compression?</em> puts that question in its title. And token deletion is the sole method compared here, while <em>Context Compression Is Not One Thing</em> treats compression as more than one kind of operation."
    ],
    "limitations": "Every result here is static: compress once, read once, score. Selective retrieval, section-level pruning and in-place updates are argued from the format and walked through in a worked example; none was benchmarked over a multi-turn session, and the paper names that as its most important gap. The baseline is LLMLingua-2 alone, and the match is approximate: TE's mean ratio was 0.585 against LLMLingua-2's 50% retention, and the roughly 11-point lead comes from the more aggressive 33% setting. On key facts the two methods sit within a tenth of a point on two of three models, with no error bars and a single randomisation of distractor placement. GPT-4.1 wrote the questions and also answered them, every reader model is a closed OpenAI model, the task is multiple-choice and English-only, and one compressor (o4-mini) did all the rewriting, so how TE quality depends on the compressor is unmapped. Compression itself costs an LLM call per chunk, which pays off for reused text and not for inputs read once and thrown away.",
    "concept_terms": [
      "prompt compression",
      "token deletion",
      "semantic chunking",
      "structured symbolic rewriting",
      "LLMLingua-2",
      "compress-once, manage-continuously"
    ],
    "references": [
      "Pan et al. (2024). LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression.",
      "Jiang et al. (2023). Llmlingua: Compressing prompts for accelerated inference of large language models.",
      "Bai et al. (2024). Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks.",
      "Chevalier et al. (2023). Adapting language models to compress contexts."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/telegraph-english.txt",
    "verification": {
      "claims": [
        {
          "claim": "claim_title / meta_description: rewriting loses fewer fine-detail answers than token deletion (fine_facts drops: GPT-4o TE -3.1 vs LLMLingua-2-50 -6.3 pp; GPT-4o-mini -9.5 vs -11.8 pp)",
          "verdict": "CONFIRMED",
          "source_quote": "TE preserves more than LLMLingua-2 at matched retention"
        },
        {
          "claim": "LLMLingua-2-style compressors delete tokens a classifier scores as low-importance",
          "verdict": "CONFIRMED",
          "source_quote": "train a classifier to delete low-importance"
        },
        {
          "claim": "Token-deletion output has no structure",
          "verdict": "CONFIRMED",
          "source_quote": "token deletion produces no structure"
        },
        {
          "claim": "Dropped connectives leave the reader model to guess how fragments relate",
          "verdict": "CONFIRMED",
          "source_quote": "the downstream model has to guess the relationship"
        },
        {
          "claim": "A lone number can look dispensable until a question turns on it",
          "verdict": "CONFIRMED",
          "source_quote": "may look dispensable next to surrounding prose"
        },
        {
          "claim": "Johnson example: 27.5% increase in early detection, false positives cut by about 12%",
          "verdict": "CONFIRMED",
          "source_quote": "resulted in a 27.5% increase in early detection rates while simultaneously reducing false positives by approximately 12%"
        },
        {
          "claim": "TE rewrite of the Johnson sentence as quoted in summary_html",
          "verdict": "CONFIRMED",
          "source_quote": "ML → MEDICAL-DIAGNOSTICS: EARLY-DETECTION+27.5% ∧ FALSE-POSITIVE-12% [JOHNSON:2023]"
        },
        {
          "claim": "Sixty-eight tokens become fourteen",
          "verdict": "CONFIRMED",
          "source_quote": "Sixty-eight tokens become fourteen"
        },
        {
          "claim": "Cause, both figures and the citation each keep a slot of their own",
          "verdict": "CONFIRMED",
          "source_quote": "The causal relationship, both quantitative claims, and the citation are each on record as separate, addressable units"
        },
        {
          "claim": "430-line grammar doubles as the compressor's system prompt",
          "verdict": "CONFIRMED",
          "source_quote": "lives in a 430-line specification document that doubles as the system prompt"
        },
        {
          "claim": "One claim per line",
          "verdict": "CONFIRMED",
          "source_quote": "each line contains exactly one claim, step, event, or question"
        },
        {
          "claim": "About 40 logical and relational symbols",
          "verdict": "CONFIRMED",
          "source_quote": "substitutes verbose phrases with ∼40 logical and relational symbols"
        },
        {
          "claim": "Symbol meanings: arrow = causation, therefore-sign = conclusion, VS = contrast",
          "verdict": "CONFIRMED",
          "source_quote": "Contrast (never causal)"
        },
        {
          "claim": "Tags for scope and modality",
          "verdict": "CONFIRMED",
          "source_quote": "modality ( LIKELY: , POSSIBLE: , CONF=0.87 )"
        },
        {
          "claim": "Fidelity ranked above brevity",
          "verdict": "CONFIRMED",
          "source_quote": "fidelity over brevity—no information may be dropped unless inferable from what remains"
        },
        {
          "claim": "Compression ratio follows information density",
          "verdict": "CONFIRMED",
          "source_quote": "lets the compression ratio adapt to each document’s information density"
        },
        {
          "claim": "Compressed text doubles as a chunked index because each line is a self-contained fact",
          "verdict": "CONFIRMED",
          "source_quote": "the compressed representation is simultaneously a semantic index"
        },
        {
          "claim": "339 LongBench-v2 documents split into 4,081 chunks",
          "verdict": "CONFIRMED",
          "source_quote": "which leaves 339 documents"
        },
        {
          "claim": "4,081 chunk-level evaluation units",
          "verdict": "CONFIRMED",
          "source_quote": "producing 4,081 chunk-level evaluation units"
        },
        {
          "claim": "Compressor: o4-mini running the v5 grammar prompt",
          "verdict": "CONFIRMED",
          "source_quote": "compressed into TE using the v5 grammar prompt with OpenAI’s o4-mini model"
        },
        {
          "claim": "LLMLingua-2 at 50% retention, 33% as the more aggressive setting",
          "verdict": "CONFIRMED",
          "source_quote": "at two retention rates: 0.50 (50% kept) and 0.33 (33% kept)"
        },
        {
          "claim": "Four-option questions; GPT-4.1 wrote the questions and distractors",
          "verdict": "CONFIRMED",
          "source_quote": "GPT-4.1 also generates the QA pairs and distractors"
        },
        {
          "claim": "4,081 key-fact questions",
          "verdict": "CONFIRMED",
          "source_quote": "key_facts (4,081 QA pairs)"
        },
        {
          "claim": "801 adversarial fine-detail questions",
          "verdict": "CONFIRMED",
          "source_quote": "fine_facts (801 QA pairs) is adversarially designed"
        },
        {
          "claim": "Reader models GPT-4.1, GPT-4o, GPT-4o-mini, GPT-4.1-nano; GPT-4o-mini is the one reader run on both suites",
          "verdict": "CONFIRMED",
          "source_quote": "key_facts is run on GPT-4.1, GPT-4o-mini, and GPT-4.1-nano; fine_facts on GPT-4o and GPT-4o-mini"
        },
        {
          "claim": "Mean TE compression ratio 0.585, a 41.5% reduction (tldr: 'about two fifths')",
          "verdict": "CONFIRMED",
          "source_quote": "The mean compression ratio is 0.585—a 41.5% token reduction"
        },
        {
          "claim": "Ratio range 0.13 to 1.57 (as printed in the prose; Table 7 lists the minimum as 0.000)",
          "verdict": "CONFIRMED",
          "source_quote": "with a range from 0.13 to 1.57"
        },
        {
          "claim": "Verbose narrative shrank hard",
          "verdict": "CONFIRMED",
          "source_quote": "verbose narrative text yields ratios of 5:1 or better"
        },
        {
          "claim": "A few short, dense chunks came out longer, because the grammar will not drop information",
          "verdict": "CONFIRMED",
          "source_quote": "very short inputs that are already informationally dense occasionally expand under TE"
        },
        {
          "claim": "GPT-4.1 scored 99.1% on TE text (key facts; original 1.000, LLMLingua-2-50 0.990); stayed within a point of original",
          "verdict": "CONFIRMED",
          "source_quote": "TE preserves 99.1% accuracy on key facts with GPT-4.1"
        },
        {
          "claim": "TE and LLMLingua-2-50 within a tenth of a point on GPT-4.1 and GPT-4.1-nano (0.991 vs 0.990; 0.950 vs 0.949)",
          "verdict": "CONFIRMED",
          "source_quote": "TE matches or edges out LLMLingua-2 across the board"
        },
        {
          "claim": "TE edged ahead on GPT-4o-mini key facts (0.957 vs 0.946)",
          "verdict": "CONFIRMED",
          "source_quote": "1.1 pp on GPT-4o-mini"
        },
        {
          "claim": "TE at or above LLMLingua-2-50 for every reader model tested; thin on key facts, clearer on fine details",
          "verdict": "CONFIRMED",
          "source_quote": "outperforms LLMLingua-2 at matched compression ratios on every model and task tested"
        },
        {
          "claim": "Fine-detail questions cost both methods two and a half to three times as much accuracy on GPT-4o-mini (TE 9.5/3.4 = 2.8; LLMLingua-2 11.8/4.5 = 2.6)",
          "verdict": "CONFIRMED",
          "source_quote": "Compression loss runs 3–4× higher than on key facts"
        },
        {
          "claim": "GPT-4o fine details: TE dropped 3.1 pp vs 6.3 pp for LLMLingua-2-50, about half as many points (6.3 - 3.1 = 3.2)",
          "verdict": "CONFIRMED",
          "source_quote": "TE holds an advantage of 3.2 pp over LLMLingua-2 on GPT-4o"
        },
        {
          "claim": "TE lead of roughly 11 pp over LLMLingua-2 at 33% retention on GPT-4o-mini fine details",
          "verdict": "CONFIRMED",
          "source_quote": "TE’s lead grows to roughly 11 pp on GPT-4o-mini"
        },
        {
          "claim": "LLMLingua-2 at 33% fell 21 pp from baseline",
          "verdict": "CONFIRMED",
          "source_quote": "21 pp from baseline"
        },
        {
          "claim": "187 of 4,081 key-fact items (4.6%) correct on original, wrong on TE, GPT-4.1-nano",
          "verdict": "CONFIRMED",
          "source_quote": "187 (4.6%) were correct on the original and incorrect on TE for GPT-4.1-nano"
        },
        {
          "claim": "TE failures cluster on dates, units, qualifiers and numerical relationships",
          "verdict": "CONFIRMED",
          "source_quote": "Failures cluster around fine details: dates, units, conditional qualifications, and numerical relationships"
        },
        {
          "claim": "Legal chunk: 'no later than 30 calendar days' became 30D; question asked calendar vs business days",
          "verdict": "CONFIRMED",
          "source_quote": "The 30D abbreviation does not distinguish"
        },
        {
          "claim": "Deletion's losses ran largest on smaller, cheaper readers (LLMLingua-2-50 drops: GPT-4.1 1.0 pp vs GPT-4o-mini 4.5 and GPT-4.1-nano 3.1 pp on key facts; 11.8 pp on GPT-4o-mini fine details)",
          "verdict": "CONFIRMED",
          "source_quote": "smaller models are precisely the ones deployed in cost-sensitive production pipelines"
        },
        {
          "claim": "TE moves reconstruction of relationships to compression time",
          "verdict": "CONFIRMED",
          "source_quote": "offloading that reconstruction work to the compression stage"
        },
        {
          "claim": "The grammar requires numbers to stay attached to their units (design rule)",
          "verdict": "CONFIRMED",
          "source_quote": "always attached to their units"
        },
        {
          "claim": "Context assembly can keep lines, collapse a section to its heading, or drop it, without a model call",
          "verdict": "CONFIRMED",
          "source_quote": "This is cheap (string manipulation, no LLM calls)"
        },
        {
          "claim": "One expensive rewrite, then cheap edits",
          "verdict": "CONFIRMED",
          "source_quote": "one expensive LLM rewrite per document, then indefinite cheap manipulation"
        },
        {
          "claim": "Compress-once, manage-continuously is offered as an architectural argument",
          "verdict": "CONFIRMED",
          "source_quote": "an architectural argument rather than an empirical result"
        },
        {
          "claim": "Every result is static: compress once, read once, score",
          "verdict": "CONFIRMED",
          "source_quote": "compress once, read once, evaluate"
        },
        {
          "claim": "Selective retrieval, pruning and updates walked through in a worked example; not benchmarked over a multi-turn session; the paper's most important gap",
          "verdict": "CONFIRMED",
          "source_quote": "This is the most important gap in the current evaluation"
        },
        {
          "claim": "LLMLingua-2 is the sole baseline / token deletion the sole method compared",
          "verdict": "CONFIRMED",
          "source_quote": "We benchmark against LLMLingua-2 only"
        },
        {
          "claim": "Approximate match: TE mean ratio 0.585 (58.5% of tokens kept) vs LLMLingua-2 50% retention",
          "verdict": "CONFIRMED",
          "source_quote": "TE’s mean compression ratio of 0.585"
        },
        {
          "claim": "No error bars; a single randomisation of distractor placement",
          "verdict": "CONFIRMED",
          "source_quote": "We did not run multiple seeds of distractor randomization"
        },
        {
          "claim": "Every reader model is a closed OpenAI model",
          "verdict": "CONFIRMED",
          "source_quote": "Our benchmark relies on OpenAI models that are not open-weight"
        },
        {
          "claim": "Multiple-choice task measuring whether a reader can still answer",
          "verdict": "CONFIRMED",
          "source_quote": "We design a multiple-choice protocol that isolates comprehension"
        },
        {
          "claim": "English-only",
          "verdict": "CONFIRMED",
          "source_quote": "English only. The grammar and benchmarks are English"
        },
        {
          "claim": "Dependence of TE quality on the compressor model is unmapped",
          "verdict": "CONFIRMED",
          "source_quote": "we have not yet mapped this sensitivity curve"
        },
        {
          "claim": "An LLM call per chunk pays off for reused text, not for inputs read once",
          "verdict": "CONFIRMED",
          "source_quote": "TE is poorly suited for compressing ephemeral inputs that will be read once and discarded"
        },
        {
          "claim": "Reference: Pan et al. (2024) LLMLingua-2",
          "verdict": "CONFIRMED",
          "source_quote": "LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression"
        },
        {
          "claim": "Reference: Jiang et al. (2023) LLMLingua",
          "verdict": "CONFIRMED",
          "source_quote": "Llmlingua: Compressing prompts for accelerated inference of large language models"
        },
        {
          "claim": "Reference: Bai et al. (2024) LongBench v2",
          "verdict": "CONFIRMED",
          "source_quote": "Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks"
        },
        {
          "claim": "Reference: Chevalier et al. (2023)",
          "verdict": "CONFIRMED",
          "source_quote": "Adapting language models to compress contexts"
        },
        {
          "claim": "Concept term: structured symbolic rewriting (also present: prompt compression, token deletion, semantic chunking, LLMLingua-2, compress-once, manage-continuously)",
          "verdict": "CONFIRMED",
          "source_quote": "Semantic Prompt Compression via Structured Symbolic Rewriting"
        },
        {
          "claim": "CORRECTED (summary_html[0]): 'asks an LLM for four things' was a count attached to a writer-defined list; the paper's four principles are a different set (fidelity, atomic lines, upper-case default, ~5x target). Count removed.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (summary_html[1]): 'nearly three times as much accuracy' on GPT-4o-mini; the table cells give 9.5/3.4 = 2.8 for TE and 11.8/4.5 = 2.6 for LLMLingua-2. Reworded to 'two and a half to three times'.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (editorial_html[0]): 'numbers stay attached to their units' was stated as an outcome; the source states it as a design rule and its error analysis finds failures on units. Reworded as a grammar requirement.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "REMOVED (editorial_html[2]): 'Context Compression Is Not One Thing pits it against a coherent summary at matched budget on multi-hop QA'; not in the source text.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "REMOVED (editorial_html[2]): 'TE is the representation the later Telegraph English papers keep returning to' and the content descriptions of the other two linked papers; not in the source text. References reduced to what each title states.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}