Two LLM encoders hit a requested token budget on 33 of 9,600 rewrites; the median ran long
Evaluating Relational Context Compression at Realized Token Budgets
Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets
ICLR 2027 submission, September 2026
Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths.
What we did and found
“Orin’s adviser is Sana, and Sana works at North Lab” can become Orin | adviser | Sana; Sana | workplace | North Lab; the shared name joins the two relations a reader needs to say where Orin’s adviser works. That rewrite is Telegraph English (TE), and the study follows it from a budget request to a reader’s answer. In a pre-registered matrix, five encoders rewrote 1,200 multi-hop contexts and three reader models answered from what came back. Full prose served as the common reference. Construction failures stayed in, as empty answers. Two of the encoders were also given token targets between 0.25 and 0.70 of the full-prose length, plus one revision after being told their count; to count as on budget, an output had to sit at or below its target and no more than two percent short. A second study reached TE by another route. An open compiler, distilled from a larger teacher, writes a source-linked plan, and a fixed serializer renders the longest variant that fits the token ceiling, without deleting a proposition or padding.
Most encoders stumbled at the interface, before any reader was involved. Six of seven in a wider census broke the output schema on most or all inputs; GPT-5.6-sol was the sole encoder to return every envelope correctly, and a mechanical recovery pass salvaged readable text from most of the malformed ones. Of 9,600 budgeted outputs, 33 landed in the band. Both encoders’ median lengths stayed above the request across the grid, and asked for a quarter of the context, Sonnet 5 delivered about two thirds. At natural length, which way TE moved relative to full prose depended on the reader: every TE encoder scored above full prose with the Mistral reader and below it with the Llama and Qwen readers, and the encoder ranking shifted as well. The compiler had the converse problem. Its in-domain plans were mostly valid and its text never overran the ceiling, yet its requested-budget F1 area was 0.0658 against 0.3936 for the pruning baseline, no reader family passed, and it failed its pre-registered acceptance test.
Key numbers
| Budgeted outputs inside the registered bandtwo anchor encoders at four requested ratios; band is at or below target and no more than two percent short | 33 of 9,600 |
| Mean token ratio for Sonnet 5 at a requested 0.25dataset-balanced context tokens over full prose, Qwen tokenizer, construction failures placed at the requested ratio; GPT-5.6-sol: 0.3496 at the same request | 0.6574 |
| Encoders that broke the output schema on most or all inputsgeneration census including two extra Ministral sizes; recovered text was still available for most rows | Six of seven |
| Dataset-balanced paired contrast, TE minus full prosenatural-length matrix, descriptive; 95% interval −0.0356 to −0.0146; construction failures scored as zero | −0.0251 |
| Requested-budget F1 area of the distilled compilerin-domain datasets; 0.3936 for budget-targeted LongLLMLingua, difference −0.328 (95% interval −0.349 to −0.307) | 0.0658 |
What this does not show
Every registered confirmatory comparison was unavailable: the same-encoder prose rewrite failed its development requirement, the pruning baseline had no matrix reader pass, a second baseline had unresolved usage rights, and the targeted-prompt calibration was not generated. The matrix results against full prose are therefore descriptive, with an unadjusted interval and failures scored as zero. Recovering text from a malformed envelope shows that text exists, not that it preserves the source, and a valid compiler plan passes structural checks without full span or semantic validation. The compiler was compared on the requested grid, and whether the baseline delivered equal token counts there is unverified. Its out-of-distribution component could not be scored because the 500 QASPER questions had no answers; citation-reach judging was prepared but never executed; and the second compiler changed several target properties at once, which leaves the contribution of each unresolved. Where along plan, serialization and reading the evidence was lost is not located. Scores use a modified token-overlap F1 that differs from multiset F1 when words repeat, the data are English multi-hop and research-paper QA on the named model versions, and raw contexts, predictions and compiler weights are missing from the archive, so the experiments cannot be rerun from it.
The authors' abstract
Context compression must preserve answerable information within a token budget, yet shorter text alone cannot isolate the effect of its representation. We study Telegraph English, which rewrites retrieved passages into relational statements, using a pre-registered encoder–consumer matrix and an open compiler trained by distillation. Six of seven encoders violated the registered output schema on most or all inputs. Only 33 of 9,600 budgeted outputs fell within the registered band: at or below the requested budget and no more than two percent short. These findings motivate evaluating answer F1 at realized token ratios. The matrix was collected, but all registered confirmatory comparisons were unavailable because their comparator or calibration conditions were missing or ineligible. The remaining answer-F1 comparisons with full prose are descriptive and score construction failures as empty answers. The second compiler failed its pre-registered acceptance test against a budget-targeted pruning baseline. Its out-of-distribution component was unscorable because answer keys were absent, and citation-reach judging was not reached. Context-compression comparisons require measured token budgets, explicit construction failures, and available comparator evidence.
More in Telegraph English
Cite
@misc{bei2026evaluating,
title = {Evaluating Relational Context Compression at Realized Token Budgets},
author = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Han, Yan and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},
year = {2026},
note = {ICLR 2027 submission},
url = {https://telegrapher.ai/research/relational-context-compression/}
}Builds on
- Arbuzov et al. (2026). Telegraph English: Semantic prompt compression via structured symbolic rewriting.
- Bei et al. (2026). Context compression is not one thing: Readable symbolic re-expression vs. coherent summary at matched budget.
- Jiang et al. (2024). LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression.
- Jiang et al. (2023). LLMLingua: Compressing prompts for accelerated inference of large language models.
- van Miltenburg et al. (2021). Preregistering NLP research.