telegrapher

Rewriting prompts one fact per line loses fewer fine-detail answers than deleting tokens

Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting

Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets, Lee Mosbacker

NeurIPS 2026 submission, May 2026. arXiv:2605.04426

Rewriting text so each line holds one claim, with explicit symbols for cause and contrast, cut tokens by about two fifths and lost fewer answers than LLMLingua-2's token deletion, by the widest margin on fine details.

What we did and found

A sentence reporting that Johnson and colleagues' machine-learning diagnostics raised early detection by 27.5% and cut false positives by about 12% comes out as ML → MEDICAL-DIAGNOSTICS: EARLY-DETECTION+27.5% ∧ FALSE-POSITIVE-12% [JOHNSON:2023]. Sixty-eight tokens become fourteen, and the cause, both figures and the citation each keep a slot of their own. That rewrite is Telegraph English (TE). Its 430-line grammar doubles as the compressor's system prompt and sets out what it wants from the LLM: one claim per line; a fixed vocabulary of about 40 symbols in place of verbose phrasing (→ for causation, ∴ for a conclusion, VS for contrast); tags for scope and modality; and fidelity ranked above brevity. To test it, 4,081 LongBench-v2 chunks were compressed with TE and with LLMLingua-2, and reader models answered the same four-option questions on the original and on each compressed version.

TE's output averaged 0.585 of the original token count. Verbose narrative shrank hard, while a few short, dense chunks came out longer than they went in, because the grammar will not drop information to save tokens. On key facts the two methods were close: GPT-4.1 scored 99.1% on TE text, TE and LLMLingua-2 at 50% retention landed within a tenth of a point of each other on GPT-4.1 and GPT-4.1-nano, and TE edged ahead on GPT-4o-mini. Fine-detail questions cost both methods two and a half to three times as much accuracy on GPT-4o-mini, the one reader run on both suites, and there the margin opened. GPT-4o lost about half as many points reading TE as reading LLMLingua-2's output, and against LLMLingua-2 at 33% retention TE's lead reached roughly 11 points on GPT-4o-mini. TE has failures of its own, and they cluster on dates, units and qualifiers: in one legal chunk, no later than 30 calendar days became 30D, and the question asked whether the days were calendar or business days.

Key numbers

Key-fact accuracy on TE text, GPT-4.14,081 questions; 1.000 on the original text, 0.990 with LLMLingua-2 at 50% retention99.1%
Mean TE compression ratiocompressed over original tokens across 4,081 chunks, a 41.5% reduction; range 0.13 to 1.570.585
Fine-detail accuracy drop on TE text, GPT-4o801 adversarial questions; LLMLingua-2 at 50% retention dropped −6.3 pp−3.1 pp
TE lead over LLMLingua-2 at 33% retention, fine details, GPT-4o-miniroughly; LLMLingua-2 at 33% fell 21 pp from baseline; not a matched-ratio comparison11 pp
Key-fact items correct on the original but wrong on TE, GPT-4.1-nano187 of 4,081 items; failures cluster on dates, units, qualifiers and numerical relationships4.6%

What this does not show

Every result here is static: compress once, read once, score. Selective retrieval, section-level pruning and in-place updates are argued from the format and walked through in a worked example; none was benchmarked over a multi-turn session, and the paper names that as its most important gap. The baseline is LLMLingua-2 alone, and the match is approximate: TE's mean ratio was 0.585 against LLMLingua-2's 50% retention, and the roughly 11-point lead comes from the more aggressive 33% setting. On key facts the two methods sit within a tenth of a point on two of three models, with no error bars and a single randomisation of distractor placement. GPT-4.1 wrote the questions and also answered them, every reader model is a closed OpenAI model, the task is multiple-choice and English-only, and one compressor (o4-mini) did all the rewriting, so how TE quality depends on the compressor is unmapped. Compression itself costs an LLM call per chunk, which pays off for reused text and not for inputs read once and thrown away.

Every number above was checked against the paper text.
The authors' abstract

We introduce Telegraph English (TE), a prompt-compression protocol that rewrites natural language into a symbol-rich, formally-structured dialect. Where tokendeletion methods such as LLMLingua-2 train a classifier to delete low-importance tokens at a fixed ratio, TE performs a full semantic rewrite: it decomposes the input into atomic fact lines, substitutes verbose phrases with ∼40 logical and relational symbols, and lets the compression ratio adapt to each document’s information density. A consequence of the line-structure rule is that compression and semantic chunking become the same operation—each output line is an independently addressable fact, so the compressed representation is simultaneously a semantic index. We evaluate TE on 4,081 question-answer pairs from LongBench-v2 across five OpenAI models and two difficulty levels. At roughly 50% token reduction, TE preserves 99.1% accuracy on key facts with GPT-4.1 and outperforms LLMLingua-2 at matched compression ratios on every model and task tested. The gap widens on smaller models—up to 11 percentage points on fine-detail tasks—suggesting that explicit relational structure compensates for limited model capacity. We release the grammar specification, compression prompt, benchmark data, and reference implementation.

Figures

Figure 1 of the paper.
Figure 1 of the paper.
Figure 2 of the paper.
Figure 2 of the paper.
On the blogRewriting a prompt loses fewer details than deleting its tokensTelegraph English rewrites text one claim per line with explicit symbols. Against LLMLingua-2 it lost fewer answers, most visibly on fine-detail questions.

More in Telegraph English

Cite

@misc{arbuzov2026telegraph,
  title         = {Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey and Mosbacker, Lee},
  year          = {2026},
  note          = {NeurIPS 2026 submission},
  eprint        = {2605.04426},
  archivePrefix = {arXiv},
  url           = {https://telegrapher.ai/research/telegraph-english/}
}

Builds on

  1. Pan et al. (2024). LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression.
  2. Jiang et al. (2023). Llmlingua: Compressing prompts for accelerated inference of large language models.
  3. Bai et al. (2024). Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks.
  4. Chevalier et al. (2023). Adapting language models to compress contexts.