Asked for a quarter of the context, one encoder sent back two thirds
Evaluating Relational Context Compression at Realized Token BudgetsThe paper: summary, reader and PDF
Line up two context compressors at a quarter of the original length and score what a reader model answers from each. Same questions, same reader, same budget. Whether that is fair depends on what each method did with the quarter. A deletion method chooses which tokens to keep, so the request acts on its output directly. A rewriting method writes a new string, and its length depends on how the model words it. For a rewriter, the budget is an instruction.
We gave that instruction to two encoders, GPT-5.6-sol and Sonnet 5: rewrite these retrieved passages into a quarter of their full-prose token count. GPT-5.6-sol came back at about a third. Sonnet 5 came back at about two thirds. The paper follows that gap from the request, through the text a reader actually receives, to the answer it gives.
A rewrite’s length is measured, not set
“Orin’s adviser is Sana, and Sana works at North Lab” can become:
Orin | adviser | Sana; Sana | workplace | North Lab
Sana now sits at the join between two relations, and those two are exactly what a reader needs to say where Orin’s adviser works. This is Telegraph English (TE): passages rewritten as compact relational statements, written by one model (the encoder) and answered from by another (the reader). Telegraph English introduced the format as a rewrite into atomic fact lines whose compression ratio adapts to each document’s information density. In our natural-length condition, TE gets no target at all. In the budgeted condition it gets one, and the question is where it lands.
A pre-registered matrix crosses five encoders (GPT-5.6-sol, Sonnet 5, Llama 4 Maverick, Magistral Small, Ministral 8B) with three readers (Qwen3.5-9B, Llama3.1-8B-Instruct, Mistral-7B-Instruct-v0.3) on 1,200 fixed questions from MuSiQue, 2WikiMultiHopQA and HotpotQA. The original context, full prose, is the common reference. In the budgeted runs, GPT-5.6-sol and Sonnet 5 got a target between 0.25 and 0.70 of the full-prose length, plus a single chance to revise after being told their token count. The band we registered in advance: at or below the target, and no more than two percent short.
Two rules shape the rest. Length is counted on the text the reader receives, under that reader’s tokenizer, with the question and reader instructions left out. And when an encoder fails to produce usable text, the row stays in, scored as an empty answer.
Little arrived in the form or length requested
Most encoders broke the interface before any reader saw a word
Each encoder was asked to return a small envelope: a field of relational text and a list of issues. In a wider census of seven encoders, which adds two more Ministral sizes, six broke that schema on most or all inputs; GPT-5.6-sol was the sole encoder to return every envelope correctly. A mechanical recovery pass pulled readable text out of most of the malformed envelopes. Llama 4 Maverick broke the schema on every raw envelope, yet recovery still supplied nonempty text for every input. Recovery shows that text exists, not that it says what the source said.
Asked for a quarter, Sonnet 5 delivered two thirds
Of the 9,600 budgeted outputs, 33 fell inside the registered band. The band is strict on purpose (its lower edge checks whether the encoder used the allowance it was given), but the misses were mostly overshoots, not near misses from below. The overshoot was stubborn: for both encoders, the median output ran over its request at every point of the grid.
Mean context-token ratio against full prose, balanced across datasets, under the Qwen tokenizer; outputs count at their measured length and construction failures at the requested ratio:
| Requested ratio | GPT-5.6-sol | Sonnet 5 |
|---|---|---|
| 0.25 | 0.3496 | 0.6574 |
| 0.4 | 0.6339 | 0.8670 |
| 0.55 | 0.7880 | 0.9838 |
| 0.7 | 0.8805 | 1.0585 |
Start with Sonnet 5 at a requested 0.7. Its mean ratio came out above 1, so on average the rewrite was longer than the prose it was meant to shorten. The 0.25 row matters more to anyone comparing methods: there one encoder handed the reader nearly twice as much text as the other, and a comparison keyed to the request would have placed them at the same budget.
The reader decides which way the result points
At natural length, how TE fares against full prose depends on who reads. With the Mistral reader, every TE encoder scored above full prose; with the Llama and Qwen readers, every one scored below. The encoder ranking moved with the reader too. Pooled with equal weight on datasets, encoders and readers, TE trailed full prose slightly — a paired contrast of −0.0251 F1 — while its complete reader prompt, question and template included, was about a ninth shorter. A result here belongs to the encoder–reader pair, not to either model alone.
The compiler held its ceiling and lost the comparison
The second study builds TE another way. Instead of prompting a general model, it trains a compiler: a Qwen3-8B student distilled from a Qwen3-Next-80B teacher on 27,284 grouped examples. The compiler writes a plan, not prose, made of propositions tied to spans of the source and the links between them. A fixed serializer renders the plan as TE, picking the longest of a fixed set of punctuation variants that fits the token ceiling and recording a failure if none does. It neither deletes propositions nor pads.
So the compiler cannot overshoot. Nor can it spend spare budget on saying more: raising the ceiling widens the choice of renderings while the plan stays fixed. A budget sweep measures how one plan is delivered.
Plans were mostly valid on the in-domain datasets. On the long QASPER research-paper inputs fewer than half were, mostly because the plan ran into the output ceiling.
The comparison arm was LongLLMLingua, a pruning method given the same budget requests, and against it the result was plain. Measured as the area under the F1 curve across the requests, normalized for width, the compiler scored 0.0658 on the in-domain data; the baseline scored 0.3936. The baseline climbed as the budget grew while the compiler stayed low. None of the three reader families met the registered positive-contrast criterion, and the compiler failed its pre-registered acceptance test.
Measure what the reader receives
The two studies show complementary limits. Given a budget, GPT-5.6-sol and Sonnet 5 mostly ran long. The compiler never overran its ceiling and still scored far below the baseline. Together they separate asking for a short input from delivering one, and delivering one from keeping inside it the evidence the question needs.
For anyone running a compression comparison, the upshot is a record: the requested budget, the string the reader actually got, its token count under that reader’s tokenizer, whether the question and template were counted, and what happened at construction. Failures stay in as failures. And the control has to match the question; full prose tells you how a compressed input compares with the original context, not whether relational form beats prose.
This is the measurement side of the Telegraph English line. Context Compression Is Not One Thing set the format against three matched-budget baselines on the same three multi-hop datasets (character-level deletion, truncation and random subsampling) and against a coherent prose summary written by the same encoder. Comparisons of that kind rest on knowing what each arm delivered. What Survives Learned Symbolic Compression? makes a parallel point from another direction: passing a well-formedness checker is not preserving the source. The compiler’s valid-plan counts are checks of that kind; they confirm that a plan is well formed, not that it is faithful.
What is still open
Every registered confirmatory comparison in the matrix is unavailable, eleven in all. The same-encoder prose rewrite failed its development requirement, the LongLLMLingua baseline had no matrix reader pass, a second baseline had unresolved usage rights, and the targeted-prompt calibration was not generated. Everything above about TE against full prose is descriptive.
The compiler comparison ran on the requested grid, and whether the baseline delivered equal token counts at each request is unverified. Its out-of-distribution component could not be scored, because the answer file had no answers for the 500 QASPER questions; citation-reach judging was prepared but never run. The second compiler changed target grouping, prompt and schema, assembly and training lengths together, so no one change can be credited or blamed. Nor can we yet say where, on the way from plan to answer, the evidence went missing. Locating it needs the per-question plans and predictions side by side, and neither the raw predictions nor the compiler weights are in the archive. All of it is English question answering on the model versions named here.
The question the matrix was registered to answer is the one still standing: whether relational statements beat a prose rewrite from the same encoder at the same delivered length.