---
type: post
title: Rewriting a retrieved passage beats cutting it to the same token budget
date: '2026-10-09'
description: At a fixed token budget, rewriting passages as entity-relation clauses
  beat truncated and thinned text on three multi-hop datasets, and a summary on one.
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Alexey Shvets
- Dmitri Kalaev
paper: https://telegrapher.ai/research/context-compression-is-not-one-thing.md
html: https://telegrapher.ai/blog/context-compression-is-not-one-thing/
---

# Rewriting a retrieved passage beats cutting it to the same token budget

A small model gets a multi-hop question and a retrieved passage too long for its budget. The answer sits at the end of a chain: a person, the city they were born in, the state that city belongs to. Trim the passage to fit and something goes. Whether the model can still answer depends on what went, an article and a verb or the city.

Compressors for this job come in a few families. A token scorer such as LLMLingua-2 decides which tokens to keep and hands back a subsequence of the input; other methods compress into the consumer's hidden states, or fine-tune a summariser on the task. In [the paper](/research/context-compression-is-not-one-thing/) we tried a different move at a fixed budget: keep the facts, and change the form they are written in.

## Entities stay, connective prose becomes operators

Here is a passage the paper takes from MuSiQue:

> Barack Obama was born in Honolulu, Hawaii. Honolulu is the capital of the state of Hawaii. Hawaii is a state in the United States.

And its rewrite:

```
Barack Obama @born Honolulu | Honolulu @capital_of Hawaii | Hawaii @state_in United States.
```

All four entities come through character for character. What disappears is the tissue between them. "Was born in" becomes `@born`, "is the capital of the state of" becomes `@capital_of`, and a pipe closes each relation before the next one starts. A reader walking the chain from Obama to the United States finds one link per clause and almost nothing else.

We call this format Telegraph English (TE). The name and the idea of rewriting rather than deleting come from [Telegraph English](/research/telegraph-english/), which set the approach against LLMLingua-2's token deletion. Here a strong frontier model, Claude Sonnet 4.6, writes TE from one fixed, task-agnostic prompt, and Qwen-3.5-9B reads it in place of the passage. Neither model is trained for the job. The prompt targets the consumer's tokenizer, so the budget we count is the budget the consumer sees.

### Four ways to spend the same tokens

For every question we counted the tokens in TE's rewrite and spent the same number on the original passage, three blunt ways: deleting characters at a regular interval until the count matched, cutting the passage off at the end, and keeping a random sample of its tokens. Each closes off a cheap explanation. If TE's gain came from squeezing characters, or from a tail that rarely matters, or if any handful of tokens would do, one of these controls would have caught it.

The fourth way is the real rival. The same encoder writes an ordinary prose summary of the passage, and the summary is truncated to TE's token count. Encoder, consumer and budget stay put; only the form of the compressed text changes.

We ran this on MuSiQue (2,417 questions, two to four hops), 2Wiki (1,500, balanced across hop levels) and HotpotQA (1,000, all two-hop). The passages are the sets each benchmark releases with its questions: the paragraphs that support the answer, mixed in with distractors.

## Cut-down passages lost on every dataset

TE minus each comparator, in F1 points (n.s. means the 95% paired-bootstrap interval spans zero):

| TE minus | MuSiQue | 2Wiki | HotpotQA |
|---|---|---|---|
| Character-deletion control | +16.6 | +13.6 | +18.2 |
| End-truncation control | +18.8 | +15.1 | +18.6 |
| Random-subset control | +18.4 | +16.3 | +20.2 |
| Prose summary, truncated | +11.94 | +0.32 (n.s.) | +1.96 (n.s.) |
| Full passage, no budget | ≈ +5 | +1.53 (n.s.) | −2.34 |

The top three rows are the clean result. Our reading of them is that these three cuts strip surface tokens without re-expressing anything, and so lose the bridge entities a multi-hop reader needs. The rewrite carries those entities over verbatim.

The summary row is narrower. On MuSiQue, the dataset where full passages score lowest, TE beat the truncated summary by 11.94 points. On 2Wiki and HotpotQA we could not tell the two apart.

### A summary cut to budget drops entities

We went looking for the reason on ten MuSiQue rows. Left whole, at a little over one and a half times TE's budget, the summary kept 78% of TE's named entities; cut to TE's budget, it kept roughly half. So the encoder was not the weak point. The cut was. Coherent prose does earn something, though: on MuSiQue the truncated summary still beat plain end-truncation by about seven points.

Our reading, which the paper calls the density argument, is about where prose spends its tokens. Articles, auxiliaries and connectives hold a sentence together, and TE's clauses hand that space to entities instead. If that is right, then at a matched budget the quantity that counts is entities per token, and TE packs in more of them.

### Against the full passage, it depends on the dataset

The last row drops the budget. TE won on MuSiQue, could not be separated from the full passage on 2Wiki, and lost on HotpotQA. Against LLMLingua-2 at its rate-50 setting the split was the same. Read plainly, the case for TE is a case for compression under a budget; when the whole passage fits, these results do not tell you to rewrite it.

### The advantage did not grow with depth

We also pre-registered a bolder prediction: that TE's edge over the full passage would widen as questions needed more hops. It did not. Within MuSiQue and 2Wiki the hop-count slopes all leaned the predicted way, and none came near significance after false-discovery-rate correction. Across hop levels the gap stays flat or wobbles. It does not climb.

How much does that null tell us? The design was powered to catch a widening of roughly 4–5 F1 points from two hops to four; a smaller effect could still be hiding below that line. What the data rule out is the strong version, an advantage that compounds with depth. A roughly constant offset fits what we saw.

## Where the budget goes

If you fit retrieved context into a small model, the lesson concerns the form of the kept tokens, not just their count. A token-scoring compressor can drop words, but its output is still a subsequence of the input; it cannot turn "is the capital of the state of" into one operator. A summary can rephrase, but it pays prose's overhead to stay coherent. The rewrite also needs no consumer training, unlike methods that compress into a model's hidden states, and it stays readable for anyone auditing what the model was shown.

We held one strong encoder fixed so the question stayed on how the consumer reads. Whether weaker encoders can produce the format reliably is a separate, messier problem. A related paper of ours, [Evaluating Relational Context Compression at Realized Token Budgets](/research/relational-context-compression/), puts the rewrite through a pre-registered matrix of encoders and consumers and finds that most encoders broke the registered output schema. [What Survives Learned Symbolic Compression?](/research/what-survives-learned-symbolic-compression/) looks at learned translators and argues that what a compressed code states about its source has to be measured apart from what a consumer recovers from it.

## What is still open

The passages are the benchmarks' own supporting-plus-distractor sets, not the output of a deployed retriever, where distractor quality varies. The full-passage and summary results stand on Qwen-3.5-9B alone. On Mistral-7B-Instruct-v0.3 the three control gaps held on MuSiQue, at a smaller size, but long full passages did not fit in memory; the full-passage rows that did fit lean toward shorter questions, and the summary comparison was not run.

The summary comparison is matched on tokens, not on entity coverage. Its entity counts rest on ten rows, and so does the density reading, which also leaves open why the summary tied with TE on 2Wiki and HotpotQA. For the three cuts we did not count entities at all; the bridge-entity explanation there is inferred from the scores. Nor did we test TE and the summary at a larger shared budget, which would separate the advantage of the format from the advantage of compressing harder.

The metric is F1, and it is generous to partial answers. Under exact match, TE's lead over the full passage on MuSiQue shrinks to about a point, with an interval that crosses zero.

Then there is the encoder. Every rewrite in the main experiments came from Claude Sonnet 4.6 and one frozen prompt; we did not vary encoder size, family or wording. A pilot hints at more: Qwen-3.5-0.8B, fine-tuned on entity-preserving rows, beat both LLMLingua-2 and the Sonnet encoder on MuSiQue at about a tenth of TE's budget. But it was handed the gold bridge entities at inference and ran on one dataset, which makes it an upper bound rather than something to deploy. Whether a small encoder can find those entities on its own, and pack them as densely, is the experiment that comes next.
