---
type: paper
slug: context-compression-is-not-one-thing
title: 'Context Compression Is Not One Thing: Readable Symbolic Re-expression vs.
  Coherent Summary at Matched Budget'
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Alexey Shvets
- Dmitri Kalaev
date: '2026-05-25'
status: ACL ARR 2026 submission
line: Telegraph English
pages: 13
html: https://telegrapher.ai/research/context-compression-is-not-one-thing/
pdf: https://telegrapher.ai/papers/context-compression-is-not-one-thing/context-compression-is-not-one-thing.pdf
reader: https://telegrapher.ai/research/context-compression-is-not-one-thing/read/
json: https://telegrapher.ai/api/papers/context-compression-is-not-one-thing.json
arxiv: https://arxiv.org/abs/2606.14875
---

# Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

## Paper gist

- **Claim:** At the same token budget, entity-relation rewrites beat cut-down passages by 13 to 20 F1 points
- **TL;DR:** Rewriting retrieved passages as pipe-separated entity-relation clauses, entities kept verbatim, beat character deletion, truncation and random subsampling at the same token budget on three multi-hop benchmarks. At a fixed budget, the form of the kept text is not a detail.
- **Method:** Claude Sonnet 4.6, under one fixed prompt, rewrote the gold-plus-distractor passages for 2,417 MuSiQue, 1,500 2Wiki and 1,000 HotpotQA questions into TE, and Qwen-3.5-9B answered from TE, from the full passage, from LLMLingua-2 at rate-50, from three controls cut to TE's per-question token count, and from a same-encoder prose summary truncated to that count, scored by token-level F1 with paired-bootstrap confidence intervals and a pre-registered test of whether TE's advantage grows with hop count.
- **Key result:** TE lead over the three matched-budget controls: +13.6 to +20.2 F1 points; TE lead over a same-encoder prose summary at matched budget, MuSiQue: +11.94 pp; Share of TE's named entities kept by the truncated summary: 0.497
- **Why it matters:** Teams that feed retrieved passages to small models under a token budget have the most riding on this.
- **Limits:** One encoder, Claude Sonnet 4.6 with a frozen prompt, produced every rewrite in the main experiments, so how TE depends on encoder capacity, family or prompt wording is untested.
- **Status:** ACL ARR 2026 submission, May 2026
- **Read:** reader /research/context-compression-is-not-one-thing/read/, PDF /papers/context-compression-is-not-one-thing/context-compression-is-not-one-thing.pdf, arXiv:2606.14875 https://arxiv.org/abs/2606.14875

## Abstract

We study context compression for multi-hop question answering with small language models. We propose Telegraph English, a readable symbolic format that rewrites retrieved passages into structured entity-relation statements, preserving reasoning evidence at lower token cost. In controlled experiments on MuSiQue, 2Wiki, and HotpotQA, Telegraph English outperforms three matched-budget compression baselines (character-level deletion, truncation, and random subsampling) on every dataset, with gains of 13 to 20 F1 points. It also outperforms a coherent prose summary produced by the same encoder on the hardest dataset. A pre-registered depth-interaction hypothesis is null: the advantage does not grow with reasoning depth within datasets. We interpret these results as evidence that readable symbolic re-expression preserves entity content more densely than either natural language or coherent summarization at matched token budget.

## What we did and found

Given the passage “Barack Obama was born in Honolulu, Hawaii. Honolulu is the capital of the state of Hawaii. Hawaii is a state in the United States.”, the encoder writes Barack Obama @born Honolulu | Honolulu @capital_of Hawaii | Hawaii @state_in United States. All four entities come through verbatim. The connecting verbs and articles are gone, replaced by pipes and short @-prefixed operators. That re-expression is Telegraph English (TE). A frozen frontier model, Claude Sonnet 4.6, produces it from a fixed, task-agnostic prompt, and a small consumer, Qwen-3.5-9B, reads it in place of the passage with no fine-tuning. The test holds the budget fixed. For each question, three controls cut the original passage down to TE's token count (deleting characters, dropping the tail, sampling random tokens); separately, the same encoder writes a free-prose summary, which is then truncated to that count. Full passages and LLMLingua-2 at rate-50 serve as reference points, across MuSiQue, 2Wiki and HotpotQA.

All nine comparisons with the controls favour TE, by 13.6 to 20.2 F1 points, and every paired-bootstrap 95% interval sits above zero. The truncated summary is a harder opponent. It loses to TE by 11.94 points on MuSiQue, the dataset where full passages score lowest, and cannot be separated from TE on 2Wiki or HotpotQA. The summary itself was not the problem; the cut was. On a 10-row MuSiQue sample, the summary kept roughly half of TE's named entities at TE's budget, against 78% when left whole at 1.59× the budget. Uncompressed passages split the datasets: TE leads by about five points on MuSiQue, shows no reliable difference on 2Wiki and trails by 2.34 on HotpotQA. The pre-registered prediction that TE's lead over full passages would widen with hop count did not hold, since the hop-count slopes leaned the predicted way but none was significant after false-discovery-rate correction. On a second consumer, Mistral-7B-Instruct-v0.3, the three control gaps on MuSiQue stayed positive, at a smaller size.

## Key numbers

| Measure | Value |
|---|---|
| TE lead over the three matched-budget controls | +13.6 to +20.2 F1 points |
| TE lead over a same-encoder prose summary at matched budget, MuSiQue | +11.94 pp |
| Share of TE's named entities kept by the truncated summary | 0.497 |
| TE versus the full passage, HotpotQA | −2.34 pp |
| Smallest depth widening the design could rule out | 4–5 F1 points |

## Why it matters

Teams that feed retrieved passages to small models under a token budget have the most riding on this. Token-scoring compressors such as LLMLingua-2 choose which tokens to keep and return a subsequence of the input; they can drop words, but they cannot rewrite the link between two entities as an operator. Holding the budget fixed isolates the form of the kept text, and form turns out to matter a great deal. The paper reads the control results this way: cutting surface text, by characters, by the tail or at random, loses the bridge entities a multi-hop reader needs. A coherent summary cut to the same length kept roughly half of TE's named entities in a ten-row check. The paper's name for the explanation is the density argument: pipe-separated clauses spend tokens on entities instead of on the articles, auxiliaries and connectives that prose needs, so at matched budget the axis that counts is entities per token.

Two practical properties come with the format. The consumer needs no training, unlike hidden-state methods that compress into the consumer's latent space, and the compressed context stays readable text that a person can audit. Neither is a reason to rewrite passages that already fit. Against full passages TE led on MuSiQue, showed no reliable difference on 2Wiki and lost on HotpotQA; the evidence is for compression under a budget.

The name and the rewrite-instead-of-delete idea come from Telegraph English, which set the approach against LLMLingua-2's token deletion. This paper holds one strong encoder fixed. Evaluating Relational Context Compression at Realized Token Budgets puts the rewrite through a pre-registered matrix of encoders and consumers, finds that most encoders broke the registered output schema, and argues for comparing at realized token ratios. What Survives Learned Symbolic Compression? turns to learned translators and holds that what a code states about its source and what a consumer recovers from it should be measured separately, a distinction that bears directly on the small-encoder question left open here.

## What this does not show

One encoder, Claude Sonnet 4.6 with a frozen prompt, produced every rewrite in the main experiments, so how TE depends on encoder capacity, family or prompt wording is untested. A pilot in which a fine-tuned Qwen-3.5-0.8B beat both LLMLingua-2 and the Sonnet encoder used oracle bridge entities at inference, ran on MuSiQue alone and sat at roughly a tenth of TE's budget; it is an upper bound on encoder substitution, not a deployable encoder. Passages are the gold-plus-distractor sets released with each benchmark rather than the output of a deployed retriever. The full-passage and summary comparisons rest on Qwen-3.5-9B. On Mistral-7B-Instruct-v0.3 only the matched-budget controls on MuSiQue give a usable comparison: long full passages did not fit at 24 GB, so the rows that did fit skew toward shorter questions, and the summary comparison was not run. The summary was compared at matched tokens, not matched entity coverage; TE and the summary were not compared at a larger shared budget; and the entity-coverage figures come from 10 MuSiQue rows. The depth null is bounded: the design rules out widenings of about 4–5 F1 points from two hops to four and cannot see smaller ones. Under exact match, TE's MuSiQue lead over full passages shrinks to +1.08 points with an interval crossing zero, consistent with part of the F1 lead being partial credit. The between-dataset pattern rests on three datasets and is reported as exploratory.

## Blog post

[Rewriting a retrieved passage beats cutting it to the same token budget](https://telegrapher.ai/blog/context-compression-is-not-one-thing.md)

## Cite

```bibtex
@misc{bei2026context,
  title         = {Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Shvets, Alexey and Kalaev, Dmitri},
  year          = {2026},
  note          = {ACL ARR 2026 submission},
  eprint        = {2606.14875},
  archivePrefix = {arXiv},
  url           = {https://telegrapher.ai/research/context-compression-is-not-one-thing/}
}
```
