telegrapher

At the same token budget, entity-relation rewrites beat cut-down passages by 13 to 20 F1 points

Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev

ACL ARR 2026 submission, May 2026. arXiv:2606.14875

Rewriting retrieved passages as pipe-separated entity-relation clauses, entities kept verbatim, beat character deletion, truncation and random subsampling at the same token budget on three multi-hop benchmarks. At a fixed budget, the form of the kept text is not a detail.

What we did and found

Given the passage “Barack Obama was born in Honolulu, Hawaii. Honolulu is the capital of the state of Hawaii. Hawaii is a state in the United States.”, the encoder writes Barack Obama @born Honolulu | Honolulu @capital_of Hawaii | Hawaii @state_in United States. All four entities come through verbatim. The connecting verbs and articles are gone, replaced by pipes and short @-prefixed operators. That re-expression is Telegraph English (TE). A frozen frontier model, Claude Sonnet 4.6, produces it from a fixed, task-agnostic prompt, and a small consumer, Qwen-3.5-9B, reads it in place of the passage with no fine-tuning. The test holds the budget fixed. For each question, three controls cut the original passage down to TE's token count (deleting characters, dropping the tail, sampling random tokens); separately, the same encoder writes a free-prose summary, which is then truncated to that count. Full passages and LLMLingua-2 at rate-50 serve as reference points, across MuSiQue, 2Wiki and HotpotQA.

All nine comparisons with the controls favour TE, by 13.6 to 20.2 F1 points, and every paired-bootstrap 95% interval sits above zero. The truncated summary is a harder opponent. It loses to TE by 11.94 points on MuSiQue, the dataset where full passages score lowest, and cannot be separated from TE on 2Wiki or HotpotQA. The summary itself was not the problem; the cut was. On a 10-row MuSiQue sample, the summary kept roughly half of TE's named entities at TE's budget, against 78% when left whole at 1.59× the budget. Uncompressed passages split the datasets: TE leads by about five points on MuSiQue, shows no reliable difference on 2Wiki and trails by 2.34 on HotpotQA. The pre-registered prediction that TE's lead over full passages would widen with hop count did not hold, since the hop-count slopes leaned the predicted way but none was significant after false-discovery-rate correction. On a second consumer, Mistral-7B-Instruct-v0.3, the three control gaps on MuSiQue stayed positive, at a smaller size.

Key numbers

TE lead over the three matched-budget controlscharacter deletion, end-truncation and random subsampling on MuSiQue, 2Wiki and HotpotQA; all nine 95% intervals strictly positive+13.6 to +20.2 F1 points
TE lead over a same-encoder prose summary at matched budget, MuSiQuesummary truncated to TE's per-question token budget; no reliable difference on 2Wiki or HotpotQA+11.94 pp
Share of TE's named entities kept by the truncated summarymedian on a 10-row MuSiQue sample at TE's budget; 0.782 before truncation, at 1.59× the budget0.497
TE versus the full passage, HotpotQATE trails uncompressed text here; it leads on MuSiQue, and the interval spans zero on 2Wiki−2.34 pp
Smallest depth widening the design could rule outacross the hop-2-to-hop-4 range at about 80% power; the pre-registered depth interaction came out null4–5 F1 points

What this does not show

One encoder, Claude Sonnet 4.6 with a frozen prompt, produced every rewrite in the main experiments, so how TE depends on encoder capacity, family or prompt wording is untested. A pilot in which a fine-tuned Qwen-3.5-0.8B beat both LLMLingua-2 and the Sonnet encoder used oracle bridge entities at inference, ran on MuSiQue alone and sat at roughly a tenth of TE's budget; it is an upper bound on encoder substitution, not a deployable encoder. Passages are the gold-plus-distractor sets released with each benchmark rather than the output of a deployed retriever. The full-passage and summary comparisons rest on Qwen-3.5-9B. On Mistral-7B-Instruct-v0.3 only the matched-budget controls on MuSiQue give a usable comparison: long full passages did not fit at 24 GB, so the rows that did fit skew toward shorter questions, and the summary comparison was not run. The summary was compared at matched tokens, not matched entity coverage; TE and the summary were not compared at a larger shared budget; and the entity-coverage figures come from 10 MuSiQue rows. The depth null is bounded: the design rules out widenings of about 4–5 F1 points from two hops to four and cannot see smaller ones. Under exact match, TE's MuSiQue lead over full passages shrinks to +1.08 points with an interval crossing zero, consistent with part of the F1 lead being partial credit. The between-dataset pattern rests on three datasets and is reported as exploratory.

Every number above was checked against the paper text.
The authors' abstract

We study context compression for multi-hop question answering with small language models. We propose Telegraph English, a readable symbolic format that rewrites retrieved passages into structured entity-relation statements, preserving reasoning evidence at lower token cost. In controlled experiments on MuSiQue, 2Wiki, and HotpotQA, Telegraph English outperforms three matched-budget compression baselines (character-level deletion, truncation, and random subsampling) on every dataset, with gains of 13 to 20 F1 points. It also outperforms a coherent prose summary produced by the same encoder on the hardest dataset. A pre-registered depth-interaction hypothesis is null: the advantage does not grow with reasoning depth within datasets. We interpret these results as evidence that readable symbolic re-expression preserves entity content more densely than either natural language or coherent summarization at matched token budget.

On the blogRewriting a retrieved passage beats cutting it to the same token budgetAt a fixed token budget, rewriting passages as entity-relation clauses beat truncated and thinned text on three multi-hop datasets, and a summary on one.

More in Telegraph English

Cite

@misc{bei2026context,
  title         = {Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Shvets, Alexey and Kalaev, Dmitri},
  year          = {2026},
  note          = {ACL ARR 2026 submission},
  eprint        = {2606.14875},
  archivePrefix = {arXiv},
  url           = {https://telegrapher.ai/research/context-compression-is-not-one-thing/}
}

Builds on

  1. Pan et al. (2024). LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression.
  2. Jiang et al. (2023). LLMLingua: Compressing prompts for accelerated inference of large language models.
  3. Xu et al. (2024). RECOMP: Improving retrieval-augmented LMs with context compression and selective augmentation.
  4. Mu et al. (2023). Learning to compress prompts with gist tokens.
  5. Trivedi et al. (2022). MuSiQue: Multihop questions via single-hop question composition.