---
type: post
title: Rewriting a prompt loses fewer details than deleting its tokens
date: '2026-10-09'
description: Telegraph English rewrites text one claim per line with explicit symbols.
  Against LLMLingua-2 it lost fewer answers, most visibly on fine-detail questions.
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Ziwei Dong
- Dmitri Kalaev
- Alexey Shvets
- Lee Mosbacker
paper: https://telegrapher.ai/research/telegraph-english.md
html: https://telegrapher.ai/blog/telegraph-english/
---

# Rewriting a prompt loses fewer details than deleting its tokens

Somewhere in a retrieved document sits the figure 4.8%. A prompt compressor is deciding whether to keep it, and to a classifier scoring tokens for importance, a bare number can look expendable beside the prose around it. Then a question asks whether the figure was 4.8% or 4.3%. The token that went was the whole answer.

Input to a language model is billed by the token, so trimming it is the obvious lever in retrieval pipelines and agent loops. LLMLingua-2, the baseline we measured against, trains a classifier to delete low-importance tokens down to a target ratio. Headline claims, often stated more than once, survive this well. Single tokens that carry meaning fare worse: a number, a unit, a "therefore" whose removal leaves the model guessing how two fragments relate. And the output is a shorter copy of the input with no structure of its own. In [the paper](/research/telegraph-english/) we rewrote the text instead.

## A grammar that asks for one claim per line

Take this sentence:

> According to research by Johnson and colleagues (2023), the application of machine learning techniques to medical diagnostics resulted in a 27.5% increase in early detection rates while simultaneously reducing false positives by approximately 12% compared to traditional methods.

After rewriting:

```
ML → MEDICAL-DIAGNOSTICS: EARLY-DETECTION+27.5% ∧ FALSE-POSITIVE-12% [JOHNSON:2023]
```

Sixty-eight tokens became fourteen. The cause sits left of an arrow; both effects are signed quantities joined by an explicit "and", and the citation has a slot of its own. As for "the application of … resulted in", it has collapsed into the arrow.

We call this Telegraph English (TE). Its grammar runs to 430 lines and doubles as the compressor's system prompt, asking the LLM for one claim, step or event per line. The Johnson line shows how much a single claim can hold: a cause and two effects. The grammar supplies about 40 symbols, each with one meaning, such as `→` for causation, `∴` for a conclusion and `VS` for a contrast that is never causal. Tags mark scope (`CTX:`) and modality (`LIKELY:`). Over all of it sits a priority order — fidelity outranks brevity, and nothing may be dropped unless it can be inferred from what remains.

That rule makes the compression ratio an outcome, not a setting. Verbose narrative shrinks fivefold or more. Dense technical text barely moves, and a few short, dense chunks come out longer than they went in.

### The compressed text is already chunked

Because the grammar wants one fact to a line, its output arrives divided into retrievable units. Here is the opening of the paper's worked example, a clinical-trial summary:

```
H1: CLINICAL-TRIAL OUTCOMES
CTX: PHASE-III RANDOMISED CONTROLLED-TRIAL(RCT); N=2400
PRIMARY-ENDPOINT: MORTALITY ↓ 23% VS PLACEBO; p<0.001 [SMITH:2024]
SECONDARY-ENDPOINT: HOSPITALIZATION ↓ 18%; p=0.003
ADVERSE-EVENTS: NAUSEA=12% ∧ HEADACHE=8% ∧ SERIOUS=2.1%
H1: SUBGROUP-ANALYSIS
…
```

A query about adverse events can pull the `ADVERSE-EVENTS` line and its `CTX:` scope, with no sliding window to tune. A prompt builder short on budget can keep the relevant lines whole, keep just the `H1:` heading of a less relevant section and drop the rest, without calling a model again. Compression and semantic chunking become one operation. The rewrite is the expensive step, it happens once, and what follows is string manipulation.

## Fine details are where deletion gives way

We split LongBench-v2 documents into 4,081 chunks and compressed each of them three ways: with TE, using o4-mini as the compressor, and with LLMLingua-2 at two settings, keeping 50% or 33% of tokens. GPT-4.1 wrote a four-option question per chunk, paraphrasing the correct option so that string matching could not help. Reader models then answered on the original and on each compressed version. A second, adversarial suite of 801 questions went after numerical qualifiers, conditions and secondary details, with near-miss distractors of the 4.8%-or-4.3% kind.

Accuracy drop from the original, in percentage points, with LLMLingua-2 at its 50% setting:

| Reader model, questions | TE | LLMLingua-2 (50%) |
|---|---|---|
| GPT-4.1, key facts | −0.9 | −1.0 |
| GPT-4o-mini, key facts | −3.4 | −4.5 |
| GPT-4.1-nano, key facts | −3.0 | −3.1 |
| GPT-4o, fine details | −3.1 | −6.3 |
| GPT-4o-mini, fine details | −9.5 | −11.8 |

On headline facts the two methods sit close: TE is ahead in each row, by about a point at most.

Fine details change the picture. On GPT-4o-mini, the one reader run on both suites, both methods lose two and a half to three times as much accuracy as they did on key facts, and on GPT-4o TE's loss is about half of LLMLingua-2's. Push LLMLingua-2 down to 33% and, on GPT-4o-mini, it falls 21 points below the original, leaving TE roughly 11 points ahead. TE's output ran far longer than that setting's, so the gap mixes method with budget. The paper reads the fall as deletion this deep starting to remove the very tokens the questions probe. Plausible, but it rests on two settings and one reader, and the paper does not report which tokens were cut.

Smaller readers lose more than larger ones under both methods. We think capacity explains this: a small model is less able to rebuild relationships from fragments, and TE gives it those relationships ready-made. The margin over LLMLingua-2 does not grow smoothly with size, though. On key facts, GPT-4.1-nano sits as close to LLMLingua-2 as GPT-4.1 does.

### TE's own errors sit on units and qualifiers

For GPT-4.1-nano, 4.6% of key-fact questions flipped from right on the original to wrong on TE. Those chunks had been compressed a little harder than average, and the failures cluster on dates, units, conditions and numerical relationships. One is worth telling in full. A legal chunk set a deadline "no later than 30 calendar days after receipt of written notice", and TE wrote `DEADLINE=30D-AFTER-NOTICE`. The question asked whether the days were calendar or business days. `30D` cannot say. That is a hole in the symbol vocabulary, not a slip by the compressor; rewriting can drop a distinction too.

## Compression that leaves a structure behind

Deletion's up-front cost is a classifier pass, and it pays later in accuracy, most visibly on fine details. TE spends an LLM call per chunk before anything is read, then hands the reader text with its relationships already marked. We would expect that trade to favour rewriting for reference material compressed once and read many times, and deletion for input read once and thrown away.

The accuracy table hides a second difference. LLMLingua-2 compresses the initial prompt, and whatever the model generates passes on at full length. TE is a format, so a pipeline can keep working in it. In the paper's arithmetic for a hypothetical five-step agent pipeline in which every stage reads and writes TE, savings come to $24 per thousand calls, against $7 for LLMLingua-2.

## What is still open

Everything beyond static compression rests on an architectural argument and a worked example: pulling one line for a query, collapsing a section to its heading, updating a fact mid-session. None of it was benchmarked across a multi-turn session, and the paper calls that its most important gap. Replacing a line in place is string manipulation; whether anything else depended on the old value is a question the paper does not raise. [RippleKB](/research/ripplekb/) asks a version of it across linked documents: edit one source fact, then test whether a system recovers the full constructed set of source units the edit affects.

The LLMLingua-2 match is approximate. TE's output averaged 0.585 of the original tokens, while LLMLingua-2 kept 50%. Two of the three key-fact margins are a tenth of a point, too small to read anything into without error bars, which we did not report; distractor placement was randomised once. [Evaluating Relational Context Compression at Realized Token Budgets](/research/relational-context-compression/), a separate study of TE, argues that such comparisons need measured token budgets, with answers evaluated at the token ratios each method actually realized.

The rest is narrower than we would like. LLMLingua-2 is the sole baseline, so no summary was compared; [Context Compression Is Not One Thing](/research/context-compression-is-not-one-thing/) sets TE against a coherent prose summary from the same encoder, and against three matched-budget baselines, on multi-hop questions answered by small models. GPT-4.1 wrote the questions and also answered them. Every reader model is a closed OpenAI model, and the questions are multiple-choice and in English.

Multiple-choice accuracy also says little about whether the compressed text still states what its source stated. [What Survives Learned Symbolic Compression?](/research/what-survives-learned-symbolic-compression/) takes that question up for grammar-constrained symbolic codes, measuring agreement with a constructed reference on arithmetic micro-worlds. How TE's quality depends on the model doing the rewriting is unmapped, since o4-mini did all of it. The compression step itself went unpriced too, so the point at which rewriting pays for itself remains an argument from amortisation.

On static questions, rewriting lost no more than deletion did, and on fine details noticeably less. Whether one fact per line also holds up as working memory, edited and pruned across a long session, is the experiment still to run.
