---
type: paper
slug: telegraph-english
title: 'Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting'
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Ziwei Dong
- Dmitri Kalaev
- Alexey Shvets
- Lee Mosbacker
date: '2026-05-04'
status: NeurIPS 2026 submission
line: Telegraph English
pages: 18
html: https://telegrapher.ai/research/telegraph-english/
pdf: https://telegrapher.ai/papers/telegraph-english/telegraph-english.pdf
reader: https://telegrapher.ai/research/telegraph-english/read/
json: https://telegrapher.ai/api/papers/telegraph-english.json
arxiv: https://arxiv.org/abs/2605.04426
openreview: https://openreview.net/forum?id=VNQstvpecn
---

# Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting

## Paper gist

- **Claim:** Rewriting prompts one fact per line loses fewer fine-detail answers than deleting tokens
- **TL;DR:** Rewriting text so each line holds one claim, with explicit symbols for cause and contrast, cut tokens by about two fifths and lost fewer answers than LLMLingua-2's token deletion, by the widest margin on fine details.
- **Method:** 339 LongBench-v2 documents were split into 4,081 chunks, compressed with TE (o4-mini running the v5 grammar prompt) and with LLMLingua-2 at 50% retention (33% as a more aggressive setting), and tested with four-option questions written by GPT-4.1 (4,081 on key facts, 801 adversarial fine-detail items) and answered by GPT-4.1, GPT-4o, GPT-4o-mini and GPT-4.1-nano.
- **Key result:** Key-fact accuracy on TE text, GPT-4.1: 99.1%; Mean TE compression ratio: 0.585; Fine-detail accuracy drop on TE text, GPT-4o: −3.1 pp
- **Why it matters:** Teams running retrieval and agent pipelines on smaller, cheaper reader models have the most at stake.
- **Limits:** Every result here is static: compress once, read once, score.
- **Status:** NeurIPS 2026 submission, May 2026
- **Read:** reader /research/telegraph-english/read/, PDF /papers/telegraph-english/telegraph-english.pdf, arXiv:2605.04426 https://arxiv.org/abs/2605.04426

## Abstract

We introduce Telegraph English (TE), a prompt-compression protocol that rewrites natural language into a symbol-rich, formally-structured dialect. Where tokendeletion methods such as LLMLingua-2 train a classifier to delete low-importance tokens at a fixed ratio, TE performs a full semantic rewrite: it decomposes the input into atomic fact lines, substitutes verbose phrases with ∼40 logical and relational symbols, and lets the compression ratio adapt to each document’s information density. A consequence of the line-structure rule is that compression and semantic chunking become the same operation—each output line is an independently addressable fact, so the compressed representation is simultaneously a semantic index. We evaluate TE on 4,081 question-answer pairs from LongBench-v2 across five OpenAI models and two difficulty levels. At roughly 50% token reduction, TE preserves 99.1% accuracy on key facts with GPT-4.1 and outperforms LLMLingua-2 at matched compression ratios on every model and task tested. The gap widens on smaller models—up to 11 percentage points on fine-detail tasks—suggesting that explicit relational structure compensates for limited model capacity. We release the grammar specification, compression prompt, benchmark data, and reference implementation.

## What we did and found

A sentence reporting that Johnson and colleagues' machine-learning diagnostics raised early detection by 27.5% and cut false positives by about 12% comes out as ML → MEDICAL-DIAGNOSTICS: EARLY-DETECTION+27.5% ∧ FALSE-POSITIVE-12% [JOHNSON:2023]. Sixty-eight tokens become fourteen, and the cause, both figures and the citation each keep a slot of their own. That rewrite is Telegraph English (TE). Its 430-line grammar doubles as the compressor's system prompt and sets out what it wants from the LLM: one claim per line; a fixed vocabulary of about 40 symbols in place of verbose phrasing (→ for causation, ∴ for a conclusion, VS for contrast); tags for scope and modality; and fidelity ranked above brevity. To test it, 4,081 LongBench-v2 chunks were compressed with TE and with LLMLingua-2, and reader models answered the same four-option questions on the original and on each compressed version.

TE's output averaged 0.585 of the original token count. Verbose narrative shrank hard, while a few short, dense chunks came out longer than they went in, because the grammar will not drop information to save tokens. On key facts the two methods were close: GPT-4.1 scored 99.1% on TE text, TE and LLMLingua-2 at 50% retention landed within a tenth of a point of each other on GPT-4.1 and GPT-4.1-nano, and TE edged ahead on GPT-4o-mini. Fine-detail questions cost both methods two and a half to three times as much accuracy on GPT-4o-mini, the one reader run on both suites, and there the margin opened. GPT-4o lost about half as many points reading TE as reading LLMLingua-2's output, and against LLMLingua-2 at 33% retention TE's lead reached roughly 11 points on GPT-4o-mini. TE has failures of its own, and they cluster on dates, units and qualifiers: in one legal chunk, no later than 30 calendar days became 30D, and the question asked whether the days were calendar or business days.

## Key numbers

| Measure | Value |
|---|---|
| Key-fact accuracy on TE text, GPT-4.1 | 99.1% |
| Mean TE compression ratio | 0.585 |
| Fine-detail accuracy drop on TE text, GPT-4o | −3.1 pp |
| TE lead over LLMLingua-2 at 33% retention, fine details, GPT-4o-mini | 11 pp |
| Key-fact items correct on the original but wrong on TE, GPT-4.1-nano | 4.6% |

## Why it matters

Teams running retrieval and agent pipelines on smaller, cheaper reader models have the most at stake. That is where compression pays for itself, and where deletion's losses ran largest in these tests. What changes is who does the reconstruction. A deletion classifier scores tokens for importance, and a number or a qualifier can look dispensable beside the prose around it, so the reader has to rebuild relationships from whatever fragments survive. TE does that rebuilding once, at compression time: causation, contrast and conclusion arrive as explicit symbols, and the grammar requires each number to stay attached to its unit.

The bigger change is to what compressed text is. Since each TE line is one fact under a tagged heading, a context-assembly step can keep the relevant lines, collapse a section to its heading, or drop it, using string operations and no further model call. One expensive rewrite; after that, cheap edits for as long as the text stays in use. The paper offers this compress-once, manage-continuously principle as an architectural argument.

Three of this paper's loose ends point to other work in the series. The LLMLingua-2 match is approximate, with TE's mean ratio of 0.585 above the 50% retention setting, and Evaluating Relational Context Compression at Realized Token Budgets makes realized budgets the basis of comparison. Multiple-choice accuracy shows whether a reader can still pick the right answer, which tells only indirectly whether the compressed text states what its source stated; What Survives Learned Symbolic Compression? puts that question in its title. And token deletion is the sole method compared here, while Context Compression Is Not One Thing treats compression as more than one kind of operation.

## What this does not show

Every result here is static: compress once, read once, score. Selective retrieval, section-level pruning and in-place updates are argued from the format and walked through in a worked example; none was benchmarked over a multi-turn session, and the paper names that as its most important gap. The baseline is LLMLingua-2 alone, and the match is approximate: TE's mean ratio was 0.585 against LLMLingua-2's 50% retention, and the roughly 11-point lead comes from the more aggressive 33% setting. On key facts the two methods sit within a tenth of a point on two of three models, with no error bars and a single randomisation of distractor placement. GPT-4.1 wrote the questions and also answered them, every reader model is a closed OpenAI model, the task is multiple-choice and English-only, and one compressor (o4-mini) did all the rewriting, so how TE quality depends on the compressor is unmapped. Compression itself costs an LLM call per chunk, which pays off for reused text and not for inputs read once and thrown away.

## Blog post

[Rewriting a prompt loses fewer details than deleting its tokens](https://telegrapher.ai/blog/telegraph-english.md)

## Cite

```bibtex
@misc{arbuzov2026telegraph,
  title         = {Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey and Mosbacker, Lee},
  year          = {2026},
  note          = {NeurIPS 2026 submission},
  eprint        = {2605.04426},
  archivePrefix = {arXiv},
  url           = {https://telegrapher.ai/research/telegraph-english/}
}
```
