telegrapher

Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

Back to the paper page. ACL ARR 2026 submission, May 2026.

PDF

All 13 pages are shown below.

Page 1 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 1 of 13
Text of page 1
Context Compression Is Not One Thing:
Readable Symbolic Re-expression vs. Coherent Summary at Matched
Budget

Anonymous ACL submission

Abstract

We study context compression for multi-hop
question answering with small language models. We propose Telegraph English, a readable symbolic format that rewrites retrieved
passages into structured entity-relation statements, preserving reasoning evidence at lower
token cost. In controlled experiments on
MuSiQue, 2Wiki, and HotpotQA, Telegraph
English outperforms three matched-budget
compression baselines (character-level deletion, truncation, and random subsampling) on
every dataset, with gains of 13 to 20 F1 points.
It also outperforms a coherent prose summary
produced by the same encoder on the hardest
dataset. A pre-registered depth-interaction hypothesis is null: the advantage does not grow
with reasoning depth within datasets. We interpret these results as evidence that readable
symbolic re-expression preserves entity content more densely than either natural language
or coherent summarization at matched token
budget.

1

Small language models face a tension on
multi-hop question answering: retrieved naturallanguage (NL) context is expensive in tokens and
error-prone in long passages, while short retrieval
discards the bridge entities a reader needs for
multi-step reasoning. Prior context-compression
work attacks this tension through token-level scoring (Jiang et al., 2023; Pan et al., 2024), hiddenstate summarisation (Mu et al., 2023; Chevalier
et al., 2023), or task-aware abstractive summarisation (Xu et al., 2024), but each commits to selective retention or latent-space summary rather than
re-expression in surface text.
We study a different move. A small encoder
rewrites the retrieved NL passage into a readable,
rule-governed symbolic format we call Telegraph
English (TE), and a small consumer reads TE in
place of NL. TE preserves entities verbatim and

replaces connective NL tissue with pipe-separated
symbolic operators. The encoder prompt is fixed
and the encoder is frozen; the consumer needs no
fine-tuning.
Across three multi-hop benchmarks (MuSiQue,
2Wiki, HotpotQA), TE outperforms three
matched-budget controls—character-level density matching, end-truncation, and random
subsampling—on every dataset and every control,
with paired-bootstrap 95% confidence intervals
strictly positive and gains of +13.6 to +20.2 F1
points (Fig. 1). TE also outperforms a coherent
prose summary produced by the same encoder
at the same token budget on the hardest dataset
(MuSiQue, +11.94 pp). The matched-budget controls rule out character-density manipulation, NLtail dispensability, and random-token sufficiency
as alternative explanations for TE’s advantage.
We pre-registered a stronger depth-interaction
hypothesis: that TE’s advantage over NL would
grow with the reasoning depth of the question.
This hypothesis is null. All four within-dataset
interaction slopes are direction-consistent with
the prediction but none is statistically significant (FDR-corrected p > 0.41, I 2 = 0%). A
minimum-detectable- effect-size analysis bounds
the design to ruling out widenings of roughly 4–5
F1 points across the hop-2-to-hop-4 range; weaker
effects cannot be distinguished from zero at our
sample size.

Introduction

Contributions.

• A matched-budget compression mechanism
for multi-hop QA: at the same token budget,
TE beats three trivial-compression controls
and a coherent-prose summariser, supporting
the reading that TE compresses where natural
language carries redundancy.

• A pre-registered null on depth-dependent
widening, paired with a minimum-detectable-

1
Page 2 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 2 of 13
Text of page 2
effect-size analysis that bounds the scope of
the matched-budget gains to constant offsets
rather than depth-scaling advantage.

2

Context compression for language models has
emerged as a response to the growing cost of long
retrieval-augmented contexts. Our contribution
differs from prior work in mechanism: TE is a
surface-text re-expression produced by a frozen
encoder with a fixed prompt, read by a frozen consumer. We organise prior work into four families
and contrast TE with each.

output. The contrast is again mechanism rather
than ratio.

Symbolic intermediates for reasoning. Prior
work has used symbolic intermediates in the
forward direction: chain-of-thought (Wei et al.,
2022), program-aided language models (Gao et al.,
2023), structured scratchpads (Nye et al., 2021),
where the consumer produces symbolic output.
We use a symbolic intermediate in the input direction: the consumer reads symbolic context in
lieu of NL prose. To our knowledge, TE is the
first tokenizer-aware, encoder-produced, readablesymbolic context-compression baseline evaluated
as a matched-budget alternative to NL on multihop QA.

Related Work

Token-level scoring. LLMLingua (Jiang et al.,
2023) and LLMLingua-2 (Pan et al., 2024) score
individual tokens for retention or deletion via a
small model trained on information-preservation
proxies. LongLLMLingua (Jiang et al., 2024) extends scoring to document-level saliency. Our
primary baseline of this family is LLMLingua-2
at rate-50. The mechanism is selective retention:
the output is a subsequence of the input. TE’s
mechanism is re-expression: the encoder rewrites
content into an entity-preserving format in which
bridge entities are kept verbatim and connective
tissue is replaced by symbolic operators. Tokenlevel scoring cannot produce that reformatting at
matched budget.

Positioning. TE is re-expression, not retention
or latent compression. The matched- budget
controls in this paper isolate the re-expression
mechanism from three trivial alternatives, and the
coherent-prose comparator further isolates it from
generic abstractive summarisation at the same budget.

3

Method

3.1

Telegraph English

Telegraph English (TE) is a context representation
produced by an encoder language model (Claude
Sonnet 4.6 via AWS Bedrock batch inference)
with a fixed, task-agnostic prompt. The encoder
rewrites the retrieved NL passage into a sequence
of pipe-separated symbolic clauses in which entities are preserved verbatim and connective NL
tissue is replaced by short @-prefixed operators.
The prompt instructs output that is compatible
with the consumer model’s tokenizer (Qwen-3.5-
9B) so that token budget at the consumer matches
what is written. A representative pre/post pair
from MuSiQue:

Hidden-state compression. GIST (Mu et al.,
2023) compresses context into soft tokens at the
hidden-state layer. AutoCompressor (Chevalier
et al., 2023) compresses long contexts into summary vectors. CEPE (Yen et al., 2024) extends
cross-attention to a compressed summary of retrieved passages. All three operate in the consumer’s latent space and require consumer-side
training. TE operates in surface text and needs
no consumer-side training—a practical distinction
for small-model deployment where consumer-side
retraining is costly and auditability matters.

NL: “Barack Obama was born in
Honolulu, Hawaii. Honolulu
is the capital of the state of
Hawaii. Hawaii is a state in
the United States.”
TE: “Barack Obama @born
Honolulu | Honolulu @capital_of
Hawaii | Hawaii @state_in
United States.”

Task-aware abstractive summarisation. RE-
COMP (Xu et al., 2024) compresses retrieved passages via a learned abstractive summariser finetuned on the downstream QA task. CompAct
(Yoon et al., 2024) uses a task-conditioned encoder
tuned on QA supervision. The encoder’s output
is natural language and the encoder is fine-tuned
on the task. TE’s encoder is task-agnostic, runs a
frozen prompt, and produces symbolic-structural

The full encoder prompt and additional examples are in Appendix C.

2
Page 3 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 3 of 13
Text of page 3
3.2

We pre-registered (Appendix H) a primary hypothesis that TE’s advantage over NL grows with the
reasoning depth of the question. Operationally,
we model question-level correctness as a function of representation (NL, TE, or LLMLingua-2),
centered hop count, and their interaction, fit per
dataset as a binomial generalised linear model
with cluster-robust standard errors by question.
The primary test is the sign and significance of the
representation × hop_count interaction
for TE: a negative slope means TE’s edge over NL
grows with hop count. We pool per-dataset slopes
across MuSiQue and 2Wiki by random- effects
meta-analysis and apply false-discovery-rate correction across the primary interaction family. HotpotQA is excluded from the regression because all
its questions are 2-hop. Full estimating equations,
the heterogeneity branch rule, and the randomeffects specification are in Appendix A.

3.3

Primary hypothesis: depth interaction

9 confidence intervals strictly positive” summary
is post-hoc. Full per-row specifications are in Appendix D.

3.4

Coherent-prose comparator

A natural follow-up question is whether TE’s advantage holds against a coherent-prose summary
at the same budget rather than against trivial controls. We run the same encoder under a freeprose summary prompt and post-truncate each
summary to TE’s per-row token budget. This
matched-budget contrast isolates representation
format from compression ratio: encoder, consumer, and budget are held fixed; only the surface
form of the compressed passage changes.

4

Experimental Setup

4.1

Datasets and consumer model

We evaluate on three multi-hop QA benchmarks
with distinct depth profiles: MuSiQue (Trivedi
et al., 2022) (n = 2,417 questions; hops ∈
{2, 3, 4}), 2Wiki (Ho et al., 2020) (n = 1,500;
balanced 500 per hop level), and HotpotQA (Yang
et al., 2018) (n = 1,000; all 2-hop). The consumer is Qwen-3.5-9B (base model; HF eager
bf16, greedy decoding) throughout. The encoder
for TE is Claude Sonnet 4.6 via AWS Bedrock
batch inference. Retrieved passages are the goldplus-distractor contexts released with each benchmark; NL, TE, and LLMLingua-2 all read the
same passage set per question.

Auxiliary mechanism: matched-budget
controls

We separately pre-registered an auxiliary mechanism observation (Appendix H): TE performs
semantic-preserving compression only where natural language carries redundancy to strip. The
test is whether TE outperforms three trivialcompression controls at matched per-row token
budget. Each control is computed against TE’s
per-row qwen-token count and rules out a specific
alternative explanation:

4.2

• Character-density. The NL passage is
rescaled at the character level so that its
qwen-token footprint matches TE’s. Rules
out the hypothesis that TE’s gain comes from
per-row character-density manipulation.

Prompt and answer extraction

All representations share a single neutral consumer prompt (Appendix B). The answer is extracted from the consumer’s final-line output by
regex and evaluated against the gold answer with
token-level F1, with binary correctness at F 1 ≥
0.5.

• End-truncation. The NL passage is truncated from the end at the qwen-token boundary so it has TE’s per-row token count. Rules
out the hypothesis that the NL tail is dispensable.

4.3

Baselines

• NL. Full retrieved passages, unmodified. The
standard no-compression baseline.

• Random subsampling. A fixed-seed uniform subsample of qwen-token positions is
drawn from the NL passage, sized to TE’s
budget. Rules out the hypothesis that any
random subset of NL tokens would suffice.

• LLMLingua-2 at rate-50 (Pan et al., 2024).
The primary learned-token-scoring baseline.

• Three matched-budget controls (characterdensity, end-truncation, random subsampling), each sized to TE’s per-row qwentoken count. Defined in §3.3 and Appendix D.

The direction of the controls is pre-registered
(TE should beat all three); the descriptive “9 of

3
Page 4 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 4 of 13
Text of page 4
• Coherent-prose summary produced by the
same encoder and post-truncated to TE’s perrow token budget (§3.4).

4.4

We use paired-bootstrap 95% confidence intervals over questions (n boot = 10,000, seed 0) for
landmark and matched-budget contrasts. The preregistered depth-interaction tests use a binomial
GLM with cluster-robust standard errors by question, pooled across datasets by random-effects
meta-analysis, with false-discovery-rate correction across the primary interaction family (Benjamini and Hochberg, 1995; DerSimonian and
Laird, 1986). Full statistical specifications, the
pre-committed heterogeneity rule, and the sensitivity check against a random-intercept fit are in
Appendix A.

4.5

We filed a pre-registration prior to data collection. The primary depth-interaction hypothesis, the matched-budget mechanism observation,
the FDR correction family, the random-effects
pooling specification, the heterogeneity branch
rule, and the landmark kill-gate tests are all preregistered; the protocol and amendment chain are
reproduced in Appendix H.

4.6

To check that the matched-budget mechanism is
not specific to one consumer family, we re-run
the three matched-budget controls on a second
consumer, Mistral-7B-Instruct-v0.3, on MuSiQue.
Long-context NL rows do not fit at 24 GB on this
consumer, so the cross-architecture claim rests on
the matched-budget controls (where TE and the
controls have similar lengths) rather than on the
full-NL landmark; details in §6.1.

5

Results

5.1

Matched-budget mechanism: TE beats
every trivial control

Statistical protocol

Control

Dataset

TE − ctrl (pp) [95% CI]

Char-density
Char-density
Char-density
End-truncation
End-truncation
End-truncation
Random-subset
Random-subset
Random-subset

MuSiQue
2Wiki
HotpotQA
MuSiQue
2Wiki
HotpotQA
MuSiQue
2Wiki
HotpotQA

+16.6 [+14.7, +18.4]
+13.6 [+11.3, +15.9]
+18.2 [+15.2, +21.0]
+18.8 [+16.9, +20.6]
+15.1 [+12.7, +17.4]
+18.6 [+15.7, +21.6]
+18.4 [+16.5, +20.3]
+16.3 [+13.9, +18.6]
+20.2 [+17.3, +23.1]

Table 1: Matched-budget mechanism. At matched
per-row qwen-token budget, TE outperforms all three
trivial controls on all three datasets. Paired bootstrap
over questions, n boot = 10,000, seed 0.

A6: char-density
(density-matched)

A7: NL end-trunc
(tail dispensable?)

A8: random-trunc
(random subset?)

25

F1 minus control F1 (pp)

Pre-registered protocol

20

+16.6

+18.2 +18.6

+18.8 +18.4

+13.6

15

+15.1

+20.2

+16.3

10

5

0

MuSiQue

2Wiki

HotpotQA

Figure 1: Matched-budget mechanism.
TE
beats character-density, end-truncation, and randomsubsampling controls on every dataset at matched perrow token budget. Error bars: 95% paired-bootstrap
CI.

advantage. Character-density rules out per-row
character manipulation. End-truncation rules out
the hypothesis that the NL tail is dispensable. Random subsampling rules out the hypothesis that
any size-matched random subset of NL tokens
would suffice. We read the result as supporting the
pre-registered mechanism: TE compresses where
natural language carries redundancy, and trivial
alternatives that strip surface tokens but do not reexpress content lose the bridge entities a multi-hop
reader needs.

Cross-architecture replication

5.2

Landmark comparisons at full budget

Table 2 reports paired-bootstrap 95% CIs for TE
against full-budget NL and against LLMLingua-2
at rate-50. On MuSiQue, the deepest-hop dataset,
TE beats NL by +4.75 pp and LLMLingua-2 by
+7.67 pp (both p < 0.001). On 2Wiki the TE–NL
gap is +1.53 pp with a CI that narrowly spans
zero. On HotpotQA the sign flips: TE loses to NL
by −2.34 pp (p < 0.01).
The between-dataset ordering—positive significant on MuSiQue, positive null on 2Wiki, negative
significant on HotpotQA—is the empirical pattern
that motivates the post-hoc moderator analysis in

Table 1 reports paired-bootstrap 95% confidence
intervals for TE versus the three matched-budget
controls on each of MuSiQue, 2Wiki, and HotpotQA. All nine intervals exclude zero from above,
with point estimates ranging from +13.6 to +20.2
percentage points (Fig. 1).
The three controls jointly rule out three specific
alternative explanations for TE’s matched-budget

4
Page 5 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 5 of 13
Text of page 5
Dataset

TE−NL (pp) [95% CI] TE−LLMLingua-2 (pp) [95% CI]

MuSiQue +4.75 [+3.06, +6.41]
2Wiki
+1.53 [−0.17, +3.24]
HotpotQA −2.34 [−4.41, −0.34]

+7.67 [+6.00, +9.32]
+1.60 [−0.13, +3.36]
−2.89 [−4.82, −0.92]

Table 2: Landmark paired comparisons. Full-NL
and LLMLingua-2 rate-50 budget regimes. Paired
bootstrap over questions, n boot = 10,000, seed 0.
Per-dataset n: MuSiQue 2,417; 2Wiki 1,500; HotpotQA 1,000.

Dataset

TE−prose (pp) [95% CI]

MuSiQue +11.94 [+10.09, +13.72]
2Wiki
+0.32 [−1.69, +2.30]
HotpotQA
+1.96 [−0.28, +4.37]

−4.28 [−6.13, −2.43]
+1.28 [−0.77, +3.33]
−4.85 [−7.09, −2.69]

5.4

We fit the pre-registered depth-interaction model
per dataset, pool MuSiQue and 2Wiki by randomeffects meta-analysis, and apply false-discovery-rate correction across the family. Table 4 reports
the four within-dataset slopes plus their metaanalytic pool. None is statistically significant:
FDR-adjusted p-values range 0.41 to 0.92 and
the pooled meta slopes are indistinguishable from
zero (I 2 = 0% on both interaction terms, so the

NL (reference)

LLMLingua-2:hop_c est [95% CI]

p FDR

−0.064 [−0.149, +0.021]
−0.004 [−0.085, +0.077]
−0.033 [−0.091, +0.026]

0.413
0.918
0.413

(Telegraph English)

MuSiQue

F1 (%)

LLMLingua-2

2Wiki

70

50

65

45

60

40

55

35

50

45

30

2

Coherent-prose comparator at matched
budget

A natural concern with the matched-budget mechanism is whether TE’s advantage persists against
a coherent-prose summary at the same budget,
rather than against trivial controls that corrupt
surface structure. We run Claude Sonnet 4.6 as
a coherent-summary encoder on the same three
datasets and post-truncate each summary to TE’s
per-row qwen- token budget. Table 3 reports the
comparison.
On MuSiQue, the deepest-hop dataset, TE outperforms the matched-budget coherent prose by
+11.94 pp with the CI strictly positive. On 2Wiki
and HotpotQA the difference is null. Coherent
prose itself loses to full NL on MuSiQue and HotpotQA, suggesting that matched-budget truncation
of coherent summaries is a costly operation when
the budget is tight: at MuSiQue’s budget, the truncated summary retains roughly half of TE’s named
entities, while the untruncated summary at 1.59×
the budget retains 78% (Appendix J).

0.711
0.711
0.711

Table 4: Depth-interaction regression. All four withindataset slopes are direction-consistent with the preregistered negative prediction but non-significant after
FDR correction; pooled estimates are indistinguishable
from zero, I 2 = 0%.

Appendix I; we report it there because at n = 3
datasets we cannot rule out unobserved datasetconstruction factors.

5.3

p FDR

−0.018 [−0.113, +0.077]
−0.019 [−0.113, +0.075]
−0.018 [−0.085, +0.048]

MuSiQue
2Wiki
Meta

Table 3: Coherent-prose comparator at matched
per-row qwen-token budget. Paired bootstrap over
questions, n boot = 10,000, seed 0. Per-dataset n:
MuSiQue 2,417; 2Wiki 1,500; HotpotQA 1,000.

TE:hop_c est [95% CI]

MuSiQue (n=2,417)
2Wiki (n=1,500)
Meta (RE)

Scope

prose−NL (pp) [95% CI] prose−LLMLingua-2 (pp) [95% CI]

−6.71 [−8.58, −4.85]
+1.21 [−0.80, +3.23]
−4.30 [−6.60, −2.12]

Scope

3

hop count

4

40

2

3

hop count

4

Figure 2: Depth-interaction null. Per-hop F1 for NL,
TE, and LLMLingua-2 within MuSiQue and 2Wiki,
with 95% bootstrap CI shading. The TE–NL gap is flat
across hop counts, not growing with depth. HotpotQA
is excluded because all questions are 2-hop.

pre-committed heterogeneity rule did not fire).
All four point estimates are negative—directionconsistent with the pre-registered prediction—but
at these p-values the direction-consistency is indistinguishable from noise, not weak supporting evidence. Section 6.2 quantifies this with a minimumdetectable- effect-size analysis.
Figure 2 visualises the null: the TE–NL gap
is flat or non-monotone across hop levels within
both MuSiQue and 2Wiki, not rising with depth
as the pre-registered mechanism would predict.

5.5

Methodological note

Our pilot swept three consumer-prompt variants. Under one variant (role-prompted), TE’s
MuSiQue F1 varied by tens of points across seeds
relative to the default prompt because verbose
responses interact with the normalised-multiset
F1 scorer in a way that is orthogonal to context
compression. Main-body numbers use the neutral
default prompt throughout; the full prompt-by-metric table is in Appendix F.

Depth interaction is null

6

Discussion

6.1

What the matched-budget result shows

The matched-budget mechanism result is the central finding of this paper. At the same per-row

5
Page 6 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 6 of 13
Text of page 6
able. The depth-dependent effect may exist but
our sample is too small to detect it. Or the
depth-dependent mechanism may simply be absent: TE’s advantage may be a constant offset over
compressed-NL equivalents rather than a depthdependent widening over full NL.
A minimum-detectable-effect-size analysis
bounds the interpretation. The pooled meta slope
is −0.018 log-odds per hop with cluster-robust
standard error 0.034; at 80% power and α =
0.05 two-sided, the minimum detectable slope
is roughly 0.095 log-odds per hop. Translated
to F1 at our consumer’s MuSiQue baseline, that
corresponds to a TE–NL widening of about 4–5
percentage points across the hop-2-to-hop-4 range.
The design therefore rules out widenings of that
magnitude with ≈ 80% power but cannot distinguish smaller effects from no effect. The null is informative against a strong-form depth-dependent
advantage and uninformative about a weak-form
one. What our data do support is that the TE–NL
gap does not grow monotonically with hop count
within MuSiQue or 2Wiki at any appreciable magnitude (Fig. 2).

qwen-token budget, TE outperforms characterdensity, end-truncation, and random-subsampling
controls by +13.6 to +20.2 percentage points
across three datasets. The three controls exhaust
three mechanistically distinct trivial-compression
strategies—compress by character deletion, drop
the NL tail, drop random NL tokens—and none
preserves the bridge entities a multi-hop reader
needs. TE preserves bridge entities verbatim and
replaces connective NL tissue with symbolic operators, a re-expression of the same semantic content
at the same budget. We read this as evidence that
TE compresses where natural language carries redundancy: the matched-budget gain is the benefit
of that operation rather than an artefact of the compression strategies the controls implement.
The coherent-prose comparator sharpens the
reading. Coherent prose at matched budget retains
roughly half of TE’s named entities on MuSiQue,
while a longer untruncated coherent summary at
1.59× the budget retains 78%. The lost entities
are bridge entities the matched-budget truncation
drops, and that loss is a property of coherent prose
at tight budgets, not an artefact of the encoder. TE
holds entity content more densely than either NL
or coherent prose because pipe-separated triples
reserve every token for entity-bearing content that
articles, auxiliaries, and connectives would otherwise consume. We name this the density argument:
at matched budget, the relevant axis is entities-per-token, and TE’s representation is the denser one.
The matched-budget mechanism replicates on
a second consumer family. Re-running the three
controls on Mistral-7B-Instruct-v0.3 on MuSiQue
yields paired-bootstrap 95% CIs strictly positive
on all three controls (+8.62/ + 11.08/ + 10.62
pp), direction-consistent with Qwen-3.5-9B at the
smaller magnitude expected for a 7B consumer.
The full-NL landmark is not a valid baseline on
Mistral because long-context NL rows do not fit
at 24 GB and the surviving subset is biased toward
shorter (easier) questions; the cross-architecture
claim therefore rests on the matched-budget controls only.

6.2

The depth-interaction hypothesis predicted that
TE’s advantage over NL would grow with hop
count within multi-hop datasets. All four withindataset slopes came out direction-consistent but
non-significant, and the pooled estimates are indistinguishable from zero. Two readings are avail-

We tested Telegraph English, a readable symbolic
re-expression of retrieved passages, as a matchedbudget alternative to natural language for multihop question answering with small language models. Across three benchmarks, TE outperforms
three trivial-compression controls and a coherentprose summariser at the same per-row token bud-

6.3

Implications and encoder provenance

Our experiments use Claude Sonnet 4.6 as the TE
encoder, deliberately chosen as a strong frontier
model so that the consumer-reading question is
not confounded with translator capacity. A natural follow-up is whether a much smaller encoder,
given domain-matched training, could substitute.
A pilot we describe in Appendix L fine-tunes
Qwen-3.5-0.8B on 5,000 entity-preserving rows
with oracle bridge- entity conditioning, evaluates
on the held-out MuSiQue test shard, and beats
both LLMLingua-2 and the Sonnet TE encoder at
roughly 10% of the qwen-token budget. The pilot
is single-dataset and uses oracle conditioning, so
it is an upper bound on encoder substitutability
rather than a deployable alternative; the general
translator-capacity question remains open.

7

Why the null is informative

6

Conclusion
Page 7 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 7 of 13
Text of page 7
get, with all nine matched-budget confidence intervals strictly positive. A pre- registered depthinteraction hypothesis is null: the advantage does
not grow with reasoning depth, and a minimumdetectable-effect-size analysis bounds the design
to ruling out widenings of 4–5 F1 points across
hop levels. We interpret these results as evidence
that the operative property of TE is entity density
per token, and the natural follow-up is whether
smaller encoders, evaluated without oracle entity
conditioning, can preserve that density.

8

Limitations and Broader Impact

8.1

Limitations

Between-dataset gradient is n = 3. The posthoc reading of the between-dataset TE–NL gradient as correlating with dataset NL ceiling rests on
three datasets. NL-ceiling- proximity is the only
monotone candidate moderator at this n; unobserved dataset-construction factors (bridge-entity
retrievability, distractor-passage content) correlated with NL ceiling cannot be ruled out. A
confirmatory replication would require a fourth
dataset where NL ceiling varies independently of
dataset construction, and we report the gradient as
an exploratory observation only (Appendix I).

encoders.

Coherent prose loses entity coverage at
matched budget. The matched-budget coherentprose comparator truncates the encoder’s freeprose output to TE’s per-row budget. On a 10-row
MuSiQue sample this drops entity coverage from
0.782 (raw, 1.59× budget) to 0.497 (matched budget). That is a density property of coherent prose
at tight budgets, not an engineering artefact we
can remove. A reader whose deployment budget is
measured in retained entities rather than consumer
tokens should treat our prose-comparator evidence
as upper-bounded; a comparison at matched entity
coverage would allow the prose representation a
larger budget and would answer a different question.

TE-at-larger-budget counterfactual untested.
We compare TE at matched qwen-token budget to coherent prose at the same budget; we
do not test TE at 1.59× budget against coherent prose at 1.59× budget. The latter would
isolate representation- format advantage from
compression-ratio advantage and remains followup work.

Prompt-template control. A control on
MuSiQue (sub-sample n = 500) under the
role-prompted consumer prompt confirms
the methodological-artefact reading of §5.5:
NL F1 (7.61%) and TE F1 (7.14%) collapse
together under that prompt, so the collapse is
prompt-template-universal and not specific to
TE. Under the explicit-reasoning prompt on the
same sub-sample, TE F1 (34.28%) substantially
exceeds NL F1 (3.12%), suggesting TE is more
prompt- robust than raw NL under that variant.
We flag this as suggestive only because the
sub-sample is MuSiQue at n = 500.

Cross-architecture coverage is mechanism-only.
The matched-budget mechanism is tested on two
consumer families (Qwen-3.5-9B and Mistral-7B-
Instruct-v0.3) and is direction-consistent on both
(§6.1). The aggregate TE–NL landmark and the
dataset-hardness gradient, however, are tested only
on Qwen-3.5-9B; the Mistral full-NL baseline is
biased by long-context out-of-memory errors on
the NL rows, so we do not claim architecture
robustness for the landmark or for the betweendataset gradient. The coherent-prose comparator
was not run on Mistral for the same reason. Crossfamily landmark coverage at a consumer with
sufficient context-window headroom is follow-up
work.

Metric brittleness. We use token-level F1 with
the standard normalised-multiset implementation.
F1 is sensitive to response verbosity (the roleprompt artefact in §5.5 is a manifestation). We report exact-match (EM) alongside F1 on MuSiQue
in Appendix G; EM and F1 deltas track the same
direction across all five comparators, which lends
confidence that the depth- interaction null and
the matched-budget mechanism are not F1-scorer
artefacts. One divergence point is worth flagging: TE’s aggregate EM advantage over NL
on MuSiQue (+1.08 pp, 95% CI [−0.58, +2.69])
is substantially smaller than its F1 advantage

Single encoder model and frozen prompt.
TE’s encoder (Claude Sonnet 4.6 via AWS
Bedrock batch inference) and the encoder prompt
are fixed. We do not vary encoder capacity, encoder family, or prompt phrasing. A dedicated
analysis of encoder sensitivity is future work; the
present paper is about whether a frozen tokenizeraware encoder can serve as a matched-budget alternative to NL, not about the landscape of such

7
Page 8 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 8 of 13
Text of page 8
(+5.24 pp, 95% CI [+3.59, +6.89]) and the EM
CI crosses zero, consistent with the aggregate TE–
NL gap being partly carried by partial-credit tokens that F1 counts as recall and EM counts as a
miss. A full LLM-judge sensitivity sweep remains
future work.

8.2

TE is a context-compression representation that,
in our data, preserves bridge-entity spans from
NL verbatim; it is therefore no more vulnerable
to entity leakage than retrieval over the source NL
passages themselves. The encoder is frozen and
task-agnostic, so TE does not implicitly encode
downstream task supervision in its output—a property we consider favourable for auditability. We
see no application-specific harms unique to TE
relative to the NL baseline it is meant to replace.

Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara,
and Akiko Aizawa. 2020. Constructing a multi-hop
QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International
Conference on Computational Linguistics.

The depth-interaction null is bounded, not
erased. The minimum-detectable-effect-size
analysis in §6.2 shows our pooled design was
powered for per-hop interaction slopes ≳ 0.095
log-odds/hop and underpowered for slopes below
that. Readers should treat the null as informative
against strong-form depth- dependent widening
(slopes ≳ 4–5 pp across the hop-2-to-hop-4 range)
and uninformative about weak-form widening.
Closing the gap would require either substantially
larger n per dataset or a narrower prediction.

Passage-set scope. We use the gold-plus-distractor contexts released with each benchmark
as retrieval input. This isolates the representationvs-NL question from the retrieval question, but it
also means our results do not speak directly to deployed retrieval pipelines where distractor quality
varies.

Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon,
Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: Program-aided language
models. In Proceedings of the 40th International
Conference on Machine Learning.

Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing
Yang, and Lili Qiu. 2023. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language
Processing.

Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng
Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2024.
LongLLMLingua: Accelerating and enhancing llms
in long context scenarios via prompt compression.
In Proceedings of the 62nd Annual Meeting of the
Association for Computational Linguistics.

Jesse Mu, Xiang Lisa Li, and Noah D. Goodman. 2023.
Learning to compress prompts with gist tokens. In
Advances in Neural Information Processing Systems
36.

Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari,
Henryk Michalewski, Jacob Austin, David Bieber,
David Dohan, Aitor Lewkowycz, Maarten Bosma,
David Luan, Charles Sutton, and Augustus Odena.
2021. Show your work: Scratchpads for intermediate computation with language models. arXiv
preprint arXiv:2112.00114.

Broader impact

References

Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the
Royal Statistical Society: Series B (Methodological),
57(1):289–300.

Alexis Chevalier, Alexander Wettig, Anirudh Ajith,
and Danqi Chen. 2023. Adapting language models
to compress contexts. In Proceedings of the 2023
Conference on Empirical Methods in Natural Language Processing.

Rebecca DerSimonian and Nan Laird. 1986. Metaanalysis in clinical trials. Controlled Clinical Trials,
7(3):177–188.

Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin
Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor
Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao,
Lili Qiu, and Dongmei Zhang. 2024. LLMLingua-
2: Data distillation for efficient and faithful taskagnostic prompt compression. In Findings of the
Association for Computational Linguistics: ACL
2024.

Harsh Trivedi, Niranjan Balasubramanian, Tushar
Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop questions via single-hop question composition.
Transactions of the Association for Computational
Linguistics, 10.

Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten
Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le,
and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In
Advances in Neural Information Processing Systems
35.

Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024. RE-
COMP: Improving retrieval-augmented LMs with
context compression and selective augmentation. In
Proceedings of the Twelfth International Conference
on Learning Representations.

8
Page 9 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 9 of 13
Text of page 9
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and
Christopher D. Manning. 2018. HotpotQA: A
dataset for diverse, explainable multi-hop question
answering. In Proceedings of the 2018 Conference
on Empirical Methods in Natural Language Processing.

Pre-committed heterogeneity rule. Before fitting, we committed to a heterogeneity branch rule:
if I 2 > 75% on a pooled slope, pooling is not
interpretable and we fall back to per-dataset inference. The rule did not fire in our sample (I 2 = 0%
on both pooled interaction slopes).

Howard Yen, Tianyu Gao, and Danqi Chen. 2024.
Long-context language modeling with parallel context encoding. In Proceedings of the 62nd Annual
Meeting of the Association for Computational Linguistics.

Chanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong, and Jaewoo Kang. 2024. CompAct:
Compressing retrieved documents actively for question answering. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language
Processing.

Paired-bootstrap. For landmark and matchedbudget contrasts, paired-bootstrap 95% confidence
intervals resample question ids with n boot =
10,000 and seed 0, using the percentile bracket
on the bootstrap distribution.

A

The pre-registered primary hypothesis predicts
that TE’s advantage over NL grows with
reasoning depth, operationalised as a negative representation × hop_count interaction within multi-hop datasets. Let y i ∈
{0, 1} be the correctness of the consumer’s answer to question i under representation r i ∈
{NL, TE, LLMLingua-2} at hop count h i ∈
{2, 3, 4}, with y i = 1 iff token-level F1 against
the gold reference is ≥ 0.5. We fit, per dataset, a
binomial GLM with logit link

B

Consumer prompt

The default consumer prompt used for all mainbody numbers is reproduced verbatim below, with
{context} and {question} placeholders substituted at runtime:

Statistical protocol

C

Answer the question using only
the context. Return the answer
on the final line.
Context:
{context}
Question: {question}
Answer:

Telegraph English encoder prompt
and example

TE is produced by Claude Sonnet 4.6 (via managed batch inference) with a fixed prompt that
instructs entity-preserving re-expression at a target qwen-token budget. The encoder prompt is
reproduced in full in the supplementary materials;
an illustrative pre/post pair from a MuSiQue row
is:

logit P(y i = 1) = α+β r r i +β h hop c +γ r,h (r i ·hop c )
(1)
where hop c = h i − h̄ centers hop count at its
dataset mean and cluster-robust standard errors by
question_id serve as a practical proxy for the
pre-registered glmer + (1|question_id)
random-intercept specification (the two specifications agree on all inferential conclusions; sensitivity analysis in Appendix E). The pre-registered
prediction is γ TE,h < 0. HotpotQA is excluded
because all its questions are 2-hop and the withindataset slope is undefined.

NL: “Barack Obama was born in
Honolulu, Hawaii. Honolulu
is the capital of the state of
Hawaii. Hawaii is a state in
the United States.”
TE: “Barack Obama @born
Honolulu | Honolulu @capital_of
Hawaii | Hawaii @state_in
United States.”

TE preserves the entity spans (“Barack
Obama”, “Honolulu”, “Hawaii”, “United States”)
verbatim and rewrites connective NL tissue into
pipe-separated symbolic clauses with @-prefixed
operators.

Meta-analysis and multiple-comparison correction. Per-dataset interaction slopes are pooled
across MuSiQue and 2Wiki by random-effects
meta-analysis (DerSimonian–Laird; DerSimonian
and Laird, 1986), yielding a pooled point estimate, 95% Wald CI, and I 2 heterogeneity statistic.
False-discovery-rate correction (Benjamini and
Hochberg, 1995) is applied across three tests per
interaction term ({M uSiQue, 2W iki, meta}).

D

A6/A7/A8 specification

All three trivial-compression controls are operationalised per row against TE’s per-row qwentoken count. Denote by B i the qwen-token count
of TE’s passage for question i.

9
Page 10 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 10 of 13
Text of page 10
A6 (char-density). We rescale the NL passage
by a character-level density transform: drop every k-th character with k chosen per row so that
the post-transform qwen-token count matches B i .
Whitespace is preserved to keep the output humanreadable.

E

F

Wave-1a pilot numbers on MuSiQue under three consumer prompts (default,
explicit_reasoning, role_prompted)
are tabulated in Table 5 below.
Under
role_prompted, TE’s F1 varies by up
to 18.3 pp relative to default on the same rows,
driven by response-length interaction with the
normalised-multiset F1 scorer.

G

Token-level F1 with the standard normalisedmultiset implementation is sensitive to response

42.74
41.10
24.43

37.99
36.47
34.80

∆

+4.75
+4.63
−10.37

Table 5: Prompt-format sensitivity on MuSiQue (Wave-
1a pilot). The role_prompted row shows a large
negative swing for TE driven by verbose-response F1
deflation, not by an intrinsic compression loss. Mainbody numbers throughout the paper use default.

verbosity (§F). To check that the C1 ′ null and the
C2 confirmation are not F1-scorer artefacts, we
re-score the same MuSiQue predictions with exactmatch (EM) and report paired-bootstrap 95% CIs
on each representation’s EM delta versus NL at
matched budget. The MuSiQue evaluation shard
contains both F1 and EM for every row; no new
inference is required.

Cluster-robust GLM vs. glmer
sensitivity

Our primary C1 ′ specification uses a binomial GLM with cluster-robust standard
errors by question_id as a practical proxy
for the pre-registered glmer(correct
~ representation * hop_c +
random-intercept
(1|question_id))
specification. We re-fit the full glmer specification on both MuSiQue and 2Wiki as a sensitivity
analysis; the estimated representation ×
hop_c slopes agree with the cluster-robust GLM
to three decimal places on MuSiQue and two
decimal places on 2Wiki, and the inferential
conclusion (all four slopes direction-consistent,
none significant after FDR-BH) is unchanged.

default
explicit_reasoning
role_prompted

A7 (NL end-truncation). We truncate the NL
passage from its end at the qwen-token boundary
so that the truncated passage has B i qwen tokens.
Partial final words are dropped at the nearest word
boundary to avoid UTF-8 fragments.

A8 (NL random-subset). We sample B i qwentoken positions uniformly without replacement
from the NL passage (fixed seed 42) and concatenate the selected tokens with a single space between segments.
All three operate in qwen-token space, not word
or character space, so the budget matches what the
consumer LM actually sees at its tokenizer.

TE F1 (%) NL F1 (%)

Prompt

Cond

NL
TE
LLMLingua-2
A6
A7
A8

F1 mean EM mean
(%)
(%)

37.50
42.74
35.07
25.38
23.96
24.31

F1 ∆ vs NL
(pp, 95% CI)

EM ∆ vs NL
(pp, 95% CI)

28.09
—
—
29.17
+5.24 [+3.59, +6.89]
+1.08 [−0.58, +2.69]
25.78
−2.43 [−3.89, −1.00]
−2.32 [−3.68, −0.99]
17.29 −12.12 [−14.03, −10.20] −10.80 [−12.66, −8.94]
16.34 −13.54 [−15.41, −11.66] −11.75 [−13.57, −9.93]
17.34 −13.19 [−15.09, −11.30] −10.76 [−12.58, −8.94]

Table 6: F1 and EM on MuSiQue (n=2,417 questions
paired) with paired-bootstrap 95% CIs on the delta vs
NL (n boot = 10,000, seed 0). EM and F1 deltas are
direction-consistent across all five comparators. Note
that TE’s EM advantage over NL is small (+1.08 pp)
and its 95% CI crosses zero, whereas its F1 advantage (+5.24 pp) is clearly positive. The three trivialcompression controls (A6/A7/A8) show large negative
deltas on both metrics with CIs well below zero, and
LLMLingua-2 is negative on both metrics with CIs excluding zero.

The direction-consistency across F1 and EM
lends confidence that the C2 mechanism finding
(A6/A7/A8 all far below TE at matched budget)
is not an F1-scorer artefact. On the positive TE
>NL comparison, the EM CI that crosses zero is
informative: it suggests TE’s aggregate advantage
on MuSiQue is at least partly carried by partialcredit rewards where TE emits the correct bridge
entity alongside additional tokens that F1 counts
as recall but EM counts as a miss. This is consistent with the compression-where-redundancy
interpretation in §6.1 (which is about mechanism
at matched budget, not about the size of the aggregate F1 advantage). A full LLM-judge sensitivity
sweep remains future work (§8.1).

Prompt × F1-metric artefact

EM alongside F1 on MuSiQue

10
Page 11 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 11 of 13
Text of page 11
H

Pre-registration protocol and
amendments

These pre-commitments—FDR-BH correction
across {MuSiQue, 2Wiki, meta}, a heterogeneity rule (I 2 > 75%; it did not fire), an analytic
MDES bound reported alongside the null, and
explicit hypothesis-generating-only labelling of
the between-dataset moderator—jointly prevent
auxiliary-to-primary promotion after the null primary.
The pre-registration was filed prior to data collection (DOI and filing date omitted from the submission version for reviewer anonymization; restored in the camera-ready). The primary C1 ′ hypothesis, the auxiliary C2 mechanism observation
(§11.7 obs 2), the FDR-BH correction family, the
DerSimonian–Laird pooling specification, the precommitted I 2 > 75% heterogeneity rule, and the
landmark kill-gate tests K-F1-1, K-TA-1, and K-γ-
1 are all pre-registered. Two in-repo amendments
were filed during the study:

A full amendment log with SHA-stamped timestamps is included in the supplementary materials.

I

“On HotpotQA (predominantly 1–2 hop), TE
does not beat NL: ∆F 1(NL − TE) > 0. This
validates that the interaction is emergent with
depth, not a global TE-wins effect.” (preregistered document, §4)

MuSiQue

4

2

2Wiki

0

2

HotpotQA

n=3 datasets
post-hoc moderator
hypothesis-generating only

4

6

30

40

50

60

70

NL F1 (%) [dataset NL-ceiling proxy]

80

Figure 3: Post-hoc NL-ceiling-proximity observation (n = 3 datasets, hypothesis-generating only).
Between-dataset TE−NL ∆F1 (pp) plotted against
dataset NL F1 (%). With only three datasets any monotonic moderator fits similarly; we therefore show no
fitted line, no R 2 , and no coefficient. The ordering is
consistent with—but does not confirm—an NL-ceiling-proximity reading; unobserved dataset-construction
factors correlated with NL ceiling cannot be ruled out
at n = 3. Error bars: 95% paired-bootstrap CI on ∆F1.

Post-hoc observation: between-dataset
gradient

This appendix expands the brief main-body
pointer at the end of §5.4 into the full detail of
the between-dataset TE−NL gradient. We report
the gradient as an exploratory observation only,
not as a contribution.
The between-dataset ordering in Table 2
(MuSiQue +4.75 significant, 2Wiki +1.53 null,
HotpotQA −2.34 significant negative) is directionconsistent with the pre-registered P2b verbatim:

6

• §11 amendment: filed C2 (compressionwhere-redundancy, §11.7 obs 2) as a separately pre-registered auxiliary mechanism,
prior to the Wave-2 cloud run that produced
the A6/A7/A8 data reported in Table 1.

8

• §10 amendment: documented a defaultprompt bug discovered in the Wave-1a pilot
and specified the remedial Wave-1b rerun at
matched LLMLingua-2 budget.

The direction is pre-registered; the specific moderator that explains the gradient is not. In our
n = 3 datasets the ordering correlates monotonically with dataset NL F1 (MuSiQue 38.0,
2Wiki 56.0, HotpotQA 69.7). Hop-based moderators (modal hop count, mean hop depth, 4-hop
share) do not have the monotone shape required:
modal hop counts are 2/no-mode/2 respectively,
mean hop depths are 2.65/3.0/2.0, and 4-hop
shares are 17%/33%/0%. NL-ceiling-proximity
is therefore the only monotone candidate moderator at n = 3. We report it as a post-hoc
observational pattern only—unobserved datasetconstruction factors (e.g., bridge-entity retrievability, distractor-passage content) correlated with NL
ceiling cannot be ruled out. Figure 3 shows the
three points with no fitted line and an explicit
n = 3 caveat.

NL F1 (pp)

At n = 3 datasets, NL-ceiling-proximity is the
only monotone candidate moderator; unobserved
dataset-construction factors (e.g., bridge-entity retrievability, distractor-passage content) cannot be
ruled out. We therefore do not treat the gradient
as a finding of this paper. A confirmatory replication with a fourth dataset where NL ceiling varies
independently of dataset construction is the minimum prerequisite for any claim; we reserve that
for follow-up work.

Hop distribution per dataset (reported for completeness, not a candidate moderator at n =

11
Page 12 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 12 of 13
Text of page 12
3). MuSiQue: 2-hop 1,252 (51.8%), 3-hop 760
(31.4%), 4-hop 405 (16.8%); modal 2, mean 2.65.
2Wiki: 2-hop 500, 3-hop 500, 4-hop 500 (balanced); no single mode, mean 3.0. HotpotQA:
2-hop 1,000 (100%); modal 2, mean 2.0. None of
modal hop, mean hop, or 4-hop share is monotone
with the between-dataset TE − NL gradient.

J

over NL on any dataset (∆ = −6.6, −13.4, −15.7
pp on MuSiQue, 2Wiki, HotpotQA respectively);
F2-struct is cut per pre-registration. K-γ-1, a hop-
3 rescue frame intended to outperform TE on 3hop questions in a revise-and-review pass, fails:
combined MuSiQue + 2Wiki hop-3 ∆(γ − TE) =
−0.15 pp, 95% CI [−0.78, +0.48], n = 1,260. γ
is cut. Both negatives are useful: F2-struct falsifies
a “surface structure alone suffices” hypothesis,
and K-γ-1 falsifies a specific “deeper revision pass
rescues the hard hops” hypothesis.

F3 coherent-prose comparator: entity
coverage and landmark contrasts

F3 evaluates Claude Sonnet 4.6 as a frontier
coherent-summary encoder on the same three
datasets, post-truncated to each per-row TE qwentoken budget (A7-style), so F3 is evaluated at
the same budget as TE on identical consumer
rows. Entity-coverage QC on a 10-row random
MuSiQue sample (seed 0): at matched TE qwentoken budget F3 retains 0.497 of TE’s named entities, versus 0.782 for F3 raw (pre-truncation) at
1.59× TE’s budget and 0.824 for LLMLingua-2
rate-50 at 2.29× TE’s character budget (Table 7);
the drop 0.782 → 0.497 is produced entirely by
matched-budget truncation, not by the encoder.
Comparing F3 vs A7 on MuSiQue (both truncated to TE’s per-row qwen-token budget, differing only in summarization quality) gives F3−A7
≈ +7 pp: coherent summarization adds value over
raw truncation even after both lose entity coverage, so the TE−F3 contrast in Table 3 isolates the
density advantage from the summarization-quality
advantage. Paired-bootstrap 95% CIs: n boot =
10,000, seed 0, paired by question_id; perdataset n: MuSiQue 2,417; 2Wiki 1,500; HotpotQA 1,000.

Compressor

TE (Telegraph English)
F3 (coherent prose, truncated)
F3 raw (pre-truncation)
LLMLingua-2 rate-50

Entity coverage (median)

Budget regime

1.000 †
0.497
0.782
0.824

matched TE qwen-tokens
matched TE qwen-tokens
1.59× TE qwen-tokens
2.29× TE char budget

L Follow-up pilot: small-LM encoder on
MuSiQue with oracle entity
conditioning

Motivation. §6.3 chose Claude Sonnet 4.6 as
the TE encoder to isolate the consumer-reading
question from translator-capacity confounds. A
reviewer may reasonably ask whether a much
smaller encoder, given domain-matched entitypreserving training, could substitute. This pilot
is a single-dataset proof-of-concept for encoder
substitutability; it is not a full replication of the
main-experiment matrix and does not close the
general translator-capacity question.

Setup. We fine-tuned Qwen-3.5-0.8B (base) via
LoRA on 5,000 entity-preserving Wikipedia QA
rows drawn from the MuSiQue and HotpotQA
training splits (2,500 each). Each training row
pairs a question, a retrieval passage, the answer,
and the decomposition’s bridge entities (the set
of surface forms the decomposition labels as appearing in intermediate hops). The encoder is conditioned at both train and inference on the gold
bridge-entity set as a system-prompt side channel; the supervised target is a Telegraph-English
paraphrase of the passage produced by Sonnet 4.6
under a bridge-entity-aware prompt. At evaluation
we re-translate the held-out MuSiQue n=2,417
test shard through the fine-tuned encoder under the
same oracle conditioning, then send the output to
the main Qwen-3.5-9B consumer using the same
default prompt, decoder setting, and F1 scorer as
the main experiments.

Table 7: Entity-coverage QC on a 10-row random
MuSiQue sample (seed 0). † TE is the reference by
construction. F3 at matched qwen-token budget retains about half of TE’s named entities; the ∼30-pp
gap from F3 raw is produced by matched-budget truncation, not by the Sonnet encoder. LLMLingua-2’s
higher coverage is measured at a different (more permissive) budget regime.

K

Results. Evaluating all four conditions on the
same n=2,417 held-out items and pairing on question id:

Honest negatives: F2-struct and
γ-rescue

Two pre-registered kill gates fired. F2-struct (a
structured NL reformulation) does not improve

12

• E10-v2 (retrained 0.8B): F1 48.04%

• TE (Sonnet 4.6, same-job): F1 42.40%
Page 13 of Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Page 13 of 13
Text of page 13
• NL baseline: F1 37.50%

• LLMLingua-2 (rate-50, matched to TE budget): F1 35.07%

Paired bootstrap over questions (n boot = 10,000,
seed 0):

• E10-v2 − LLMLingua-2: +12.97 pp, 95%
CI [+11.00, +14.92]

• E10-v2 − TE: +5.64 pp, [+3.76, +7.54]

• E10-v2 − NL: +10.54 pp, [+8.55, +12.53]

All three CIs are strictly positive. The samejob TE F1 42.40% differs by 0.3 pp from the
main-experiment TE 42.74%, consistent with
deterministic-decode batch-order shifts across
consumer-eval runs.

binding constraint on our mechanism and density claims: a domain-matched 0.8B encoder with
oracle entity conditioning reaches consumer F1
above LLMLingua-2 and above the frontier TE
encoder at a fraction of the token budget. It does
not establish that translator capacity is unimportant across 2Wiki or HotpotQA, that small-LM
encoders match TE without oracle conditioning,
or that the density property transfers to matchedbudget generation. The general translator-capacity
question remains open and is the subject of concurrent work.

Artifacts. The following pilot artifacts accompany this paper as supplementary material: the
5,000-row entity-preserving training data; the
Sonnet-generated Telegraph-English targets; the
LoRA fine-tune config and training-loss curves;
the re-translated MuSiQue n=2,417 test shard
produced by the fine-tuned encoder; the consumereval shard, per-row F1/EM outputs, and pairedbootstrap JSON with all three contrasts.

Scope limits. Three limitations bound the reading of these numbers.
(i) Single-dataset.
The pilot tests only
MuSiQue; the main paper’s claims span three
datasets. Running the pilot on 2Wiki and HotpotQA is mechanical but was deprioritised as the
MuSiQue result already suffices to address the
“frontier-in, frontier-out” framing concern on the
dataset where TE most clearly dominates.
(ii) Budget mismatch. The retrained 0.8B output
sits at ∼ 10% of TE’s qwen-token budget (median
126 vs. 1251 per row; only 9/2417 rows exceed
TE’s budget and are truncated). The +5.64 pp
lift over TE therefore reads as an extreme-density
data point, not matched-budget dominance. A
matched-budget probe (letting E10-v2 generate
up to TE’s per-row budget) would disambiguate
extreme-density is sufficient from E10-v2 happens
to work at low budget; we do not run that probe
here.
(iii) Oracle conditioning at inference. The decomposition’s bridge-entity set is used as prompting input to the encoder at inference time. In
a deployment setting these labels are not available without either human annotation or a separate entity-extraction model; the pilot thus demonstrates an upper bound on encoder substitutability,
not a deployable alternative. A non-oracle control
(conditioning on NER-predicted bridge entities,
or on entities recovered from the consumer’s firstpass guess) is the natural follow-up.

What the pilot does and does not establish.
Subject to limits (i)–(iii), the pilot establishes
that on MuSiQue translator capacity is not the

13