---
type: paper
slug: relational-context-compression
title: Evaluating Relational Context Compression at Realized Token Budgets
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Yan Han
- Dmitri Kalaev
- Yanxin Zhang
- Alexey Shvets
date: '2026-09-18'
status: ICLR 2027 submission
line: Telegraph English
pages: 19
html: https://telegrapher.ai/research/relational-context-compression/
pdf: https://telegrapher.ai/papers/relational-context-compression/relational-context-compression.pdf
reader: https://telegrapher.ai/research/relational-context-compression/read/
json: https://telegrapher.ai/api/papers/relational-context-compression.json
openreview: https://openreview.net/forum?id=jQOAf9MQpv
---

# Evaluating Relational Context Compression at Realized Token Budgets

## Paper gist

- **Claim:** Two LLM encoders hit a requested token budget on 33 of 9,600 rewrites; the median ran long
- **TL;DR:** Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths.
- **Method:** Five LLM encoders (GPT-5.6-sol, Sonnet 5, Llama 4 Maverick, Magistral Small, Ministral 8B) rewrote 1,200 MuSiQue, 2WikiMultiHopQA and HotpotQA contexts into TE at natural length, and two of them also at requested ratios between 0.25 and 0.70, for Qwen3.5-9B, Llama3.1-8B-Instruct and Mistral-7B-Instruct-v0.3 readers with full prose as the reference; separately, a Qwen3-8B compiler distilled from Qwen3-Next-80B plans was scored against budget-targeted LongLLMLingua over the same requested grid.
- **Key result:** Budgeted outputs inside the registered band: 33 of 9,600; Mean token ratio for Sonnet 5 at a requested 0.25: 0.6574; Encoders that broke the output schema on most or all inputs: Six of seven
- **Why it matters:** Whoever is choosing between context compressors for a retrieval or agent pipeline has a stake here, more so when one candidate rewrites rather than deletes.
- **Limits:** Every registered confirmatory comparison was unavailable: the same-encoder prose rewrite failed its development requirement, the pruning baseline had no matrix reader pass, a second baseline had unresolved usage rights, and the targeted-prompt calibration was not generated.
- **Status:** ICLR 2027 submission, September 2026
- **Read:** reader /research/relational-context-compression/read/, PDF /papers/relational-context-compression/relational-context-compression.pdf

## Abstract

Context compression must preserve answerable information within a token budget, yet shorter text alone cannot isolate the effect of its representation. We study Telegraph English, which rewrites retrieved passages into relational statements, using a pre-registered encoder–consumer matrix and an open compiler trained by distillation. Six of seven encoders violated the registered output schema on most or all inputs. Only 33 of 9,600 budgeted outputs fell within the registered band: at or below the requested budget and no more than two percent short. These findings motivate evaluating answer F1 at realized token ratios. The matrix was collected, but all registered confirmatory comparisons were unavailable because their comparator or calibration conditions were missing or ineligible. The remaining answer-F1 comparisons with full prose are descriptive and score construction failures as empty answers. The second compiler failed its pre-registered acceptance test against a budget-targeted pruning baseline. Its out-of-distribution component was unscorable because answer keys were absent, and citation-reach judging was not reached. Context-compression comparisons require measured token budgets, explicit construction failures, and available comparator evidence.

## What we did and found

“Orin’s adviser is Sana, and Sana works at North Lab” can become Orin | adviser | Sana; Sana | workplace | North Lab; the shared name joins the two relations a reader needs to say where Orin’s adviser works. That rewrite is Telegraph English (TE), and the study follows it from a budget request to a reader’s answer. In a pre-registered matrix, five encoders rewrote 1,200 multi-hop contexts and three reader models answered from what came back. Full prose served as the common reference. Construction failures stayed in, as empty answers. Two of the encoders were also given token targets between 0.25 and 0.70 of the full-prose length, plus one revision after being told their count; to count as on budget, an output had to sit at or below its target and no more than two percent short. A second study reached TE by another route. An open compiler, distilled from a larger teacher, writes a source-linked plan, and a fixed serializer renders the longest variant that fits the token ceiling, without deleting a proposition or padding.

Most encoders stumbled at the interface, before any reader was involved. Six of seven in a wider census broke the output schema on most or all inputs; GPT-5.6-sol was the sole encoder to return every envelope correctly, and a mechanical recovery pass salvaged readable text from most of the malformed ones. Of 9,600 budgeted outputs, 33 landed in the band. Both encoders’ median lengths stayed above the request across the grid, and asked for a quarter of the context, Sonnet 5 delivered about two thirds. At natural length, which way TE moved relative to full prose depended on the reader: every TE encoder scored above full prose with the Mistral reader and below it with the Llama and Qwen readers, and the encoder ranking shifted as well. The compiler had the converse problem. Its in-domain plans were mostly valid and its text never overran the ceiling, yet its requested-budget F1 area was 0.0658 against 0.3936 for the pruning baseline, no reader family passed, and it failed its pre-registered acceptance test.

## Key numbers

| Measure | Value |
|---|---|
| Budgeted outputs inside the registered band | 33 of 9,600 |
| Mean token ratio for Sonnet 5 at a requested 0.25 | 0.6574 |
| Encoders that broke the output schema on most or all inputs | Six of seven |
| Dataset-balanced paired contrast, TE minus full prose | −0.0251 |
| Requested-budget F1 area of the distilled compiler | 0.0658 |

## Why it matters

Whoever is choosing between context compressors for a retrieval or agent pipeline has a stake here, more so when one candidate rewrites rather than deletes. Deletion selects tokens, so a requested ratio acts directly on what is kept. A rewriter writes a new string whose length follows its wording, which makes the request an instruction. Lined up at the same requested ratio, two rewriters can hand a reader very different amounts of text; at a requested quarter, Sonnet 5’s output was nearly twice the length of GPT-5.6-sol’s. What the paper asks for instead is a record of what was delivered: the string itself, its token count under the reader’s own tokenizer, and whether the question and template were counted.

Construction failures go in that record too. Dropping the rows where no usable text came back would change what is being scored, while keeping them as zero-F1 answers measures the whole procedure, including whether it delivered an input at all. The compiler shows the converse. A hard ceiling caps length and says nothing about content: the serializer held every ceiling it was given, and the answers still fell far below a pruning baseline.

Earlier work in the Telegraph English line compared the representation with token-matched controls and post-truncated prose summaries. This paper turns to the procedure that produces it, from request to delivered string. The controlled question of form, relations against a same-encoder prose rewrite at the same delivered length, was registered here and could not be run. It stays open.

## What this does not show

Every registered confirmatory comparison was unavailable: the same-encoder prose rewrite failed its development requirement, the pruning baseline had no matrix reader pass, a second baseline had unresolved usage rights, and the targeted-prompt calibration was not generated. The matrix results against full prose are therefore descriptive, with an unadjusted interval and failures scored as zero. Recovering text from a malformed envelope shows that text exists, not that it preserves the source, and a valid compiler plan passes structural checks without full span or semantic validation. The compiler was compared on the requested grid, and whether the baseline delivered equal token counts there is unverified. Its out-of-distribution component could not be scored because the 500 QASPER questions had no answers; citation-reach judging was prepared but never executed; and the second compiler changed several target properties at once, which leaves the contribution of each unresolved. Where along plan, serialization and reading the evidence was lost is not located. Scores use a modified token-overlap F1 that differs from multiset F1 when words repeat, the data are English multi-hop and research-paper QA on the named model versions, and raw contexts, predictions and compiler weights are missing from the archive, so the experiments cannot be rerun from it.

## Blog post

[Asked for a quarter of the context, one encoder sent back two thirds](https://telegrapher.ai/blog/relational-context-compression.md)

## Cite

```bibtex
@misc{bei2026evaluating,
  title         = {Evaluating Relational Context Compression at Realized Token Budgets},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Han, Yan and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/relational-context-compression/}
}
```
