Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget
Back to the paper page. ACL ARR 2026 submission, May 2026.
All 13 pages are shown below.
Text of page 1
Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget Anonymous ACL submission Abstract We study context compression for multi-hop question answering with small language models. We propose Telegraph English, a readable symbolic format that rewrites retrieved passages into structured entity-relation statements, preserving reasoning evidence at lower token cost. In controlled experiments on MuSiQue, 2Wiki, and HotpotQA, Telegraph English outperforms three matched-budget compression baselines (character-level deletion, truncation, and random subsampling) on every dataset, with gains of 13 to 20 F1 points. It also outperforms a coherent prose summary produced by the same encoder on the hardest dataset. A pre-registered depth-interaction hypothesis is null: the advantage does not grow with reasoning depth within datasets. We interpret these results as evidence that readable symbolic re-expression preserves entity content more densely than either natural language or coherent summarization at matched token budget. 1 Small language models face a tension on multi-hop question answering: retrieved naturallanguage (NL) context is expensive in tokens and error-prone in long passages, while short retrieval discards the bridge entities a reader needs for multi-step reasoning. Prior context-compression work attacks this tension through token-level scoring (Jiang et al., 2023; Pan et al., 2024), hiddenstate summarisation (Mu et al., 2023; Chevalier et al., 2023), or task-aware abstractive summarisation (Xu et al., 2024), but each commits to selective retention or latent-space summary rather than re-expression in surface text. We study a different move. A small encoder rewrites the retrieved NL passage into a readable, rule-governed symbolic format we call Telegraph English (TE), and a small consumer reads TE in place of NL. TE preserves entities verbatim and replaces connective NL tissue with pipe-separated symbolic operators. The encoder prompt is fixed and the encoder is frozen; the consumer needs no fine-tuning. Across three multi-hop benchmarks (MuSiQue, 2Wiki, HotpotQA), TE outperforms three matched-budget controls—character-level density matching, end-truncation, and random subsampling—on every dataset and every control, with paired-bootstrap 95% confidence intervals strictly positive and gains of +13.6 to +20.2 F1 points (Fig. 1). TE also outperforms a coherent prose summary produced by the same encoder at the same token budget on the hardest dataset (MuSiQue, +11.94 pp). The matched-budget controls rule out character-density manipulation, NLtail dispensability, and random-token sufficiency as alternative explanations for TE’s advantage. We pre-registered a stronger depth-interaction hypothesis: that TE’s advantage over NL would grow with the reasoning depth of the question. This hypothesis is null. All four within-dataset interaction slopes are direction-consistent with the prediction but none is statistically significant (FDR-corrected p > 0.41, I 2 = 0%). A minimum-detectable- effect-size analysis bounds the design to ruling out widenings of roughly 4–5 F1 points across the hop-2-to-hop-4 range; weaker effects cannot be distinguished from zero at our sample size. Introduction Contributions. • A matched-budget compression mechanism for multi-hop QA: at the same token budget, TE beats three trivial-compression controls and a coherent-prose summariser, supporting the reading that TE compresses where natural language carries redundancy. • A pre-registered null on depth-dependent widening, paired with a minimum-detectable- 1
Text of page 2
effect-size analysis that bounds the scope of the matched-budget gains to constant offsets rather than depth-scaling advantage. 2 Context compression for language models has emerged as a response to the growing cost of long retrieval-augmented contexts. Our contribution differs from prior work in mechanism: TE is a surface-text re-expression produced by a frozen encoder with a fixed prompt, read by a frozen consumer. We organise prior work into four families and contrast TE with each. output. The contrast is again mechanism rather than ratio. Symbolic intermediates for reasoning. Prior work has used symbolic intermediates in the forward direction: chain-of-thought (Wei et al., 2022), program-aided language models (Gao et al., 2023), structured scratchpads (Nye et al., 2021), where the consumer produces symbolic output. We use a symbolic intermediate in the input direction: the consumer reads symbolic context in lieu of NL prose. To our knowledge, TE is the first tokenizer-aware, encoder-produced, readablesymbolic context-compression baseline evaluated as a matched-budget alternative to NL on multihop QA. Related Work Token-level scoring. LLMLingua (Jiang et al., 2023) and LLMLingua-2 (Pan et al., 2024) score individual tokens for retention or deletion via a small model trained on information-preservation proxies. LongLLMLingua (Jiang et al., 2024) extends scoring to document-level saliency. Our primary baseline of this family is LLMLingua-2 at rate-50. The mechanism is selective retention: the output is a subsequence of the input. TE’s mechanism is re-expression: the encoder rewrites content into an entity-preserving format in which bridge entities are kept verbatim and connective tissue is replaced by symbolic operators. Tokenlevel scoring cannot produce that reformatting at matched budget. Positioning. TE is re-expression, not retention or latent compression. The matched- budget controls in this paper isolate the re-expression mechanism from three trivial alternatives, and the coherent-prose comparator further isolates it from generic abstractive summarisation at the same budget. 3 Method 3.1 Telegraph English Telegraph English (TE) is a context representation produced by an encoder language model (Claude Sonnet 4.6 via AWS Bedrock batch inference) with a fixed, task-agnostic prompt. The encoder rewrites the retrieved NL passage into a sequence of pipe-separated symbolic clauses in which entities are preserved verbatim and connective NL tissue is replaced by short @-prefixed operators. The prompt instructs output that is compatible with the consumer model’s tokenizer (Qwen-3.5- 9B) so that token budget at the consumer matches what is written. A representative pre/post pair from MuSiQue: Hidden-state compression. GIST (Mu et al., 2023) compresses context into soft tokens at the hidden-state layer. AutoCompressor (Chevalier et al., 2023) compresses long contexts into summary vectors. CEPE (Yen et al., 2024) extends cross-attention to a compressed summary of retrieved passages. All three operate in the consumer’s latent space and require consumer-side training. TE operates in surface text and needs no consumer-side training—a practical distinction for small-model deployment where consumer-side retraining is costly and auditability matters. NL: “Barack Obama was born in Honolulu, Hawaii. Honolulu is the capital of the state of Hawaii. Hawaii is a state in the United States.” TE: “Barack Obama @born Honolulu | Honolulu @capital_of Hawaii | Hawaii @state_in United States.” Task-aware abstractive summarisation. RE- COMP (Xu et al., 2024) compresses retrieved passages via a learned abstractive summariser finetuned on the downstream QA task. CompAct (Yoon et al., 2024) uses a task-conditioned encoder tuned on QA supervision. The encoder’s output is natural language and the encoder is fine-tuned on the task. TE’s encoder is task-agnostic, runs a frozen prompt, and produces symbolic-structural The full encoder prompt and additional examples are in Appendix C. 2
Text of page 3
3.2
We pre-registered (Appendix H) a primary hypothesis that TE’s advantage over NL grows with the
reasoning depth of the question. Operationally,
we model question-level correctness as a function of representation (NL, TE, or LLMLingua-2),
centered hop count, and their interaction, fit per
dataset as a binomial generalised linear model
with cluster-robust standard errors by question.
The primary test is the sign and significance of the
representation × hop_count interaction
for TE: a negative slope means TE’s edge over NL
grows with hop count. We pool per-dataset slopes
across MuSiQue and 2Wiki by random- effects
meta-analysis and apply false-discovery-rate correction across the primary interaction family. HotpotQA is excluded from the regression because all
its questions are 2-hop. Full estimating equations,
the heterogeneity branch rule, and the randomeffects specification are in Appendix A.
3.3
Primary hypothesis: depth interaction
9 confidence intervals strictly positive” summary
is post-hoc. Full per-row specifications are in Appendix D.
3.4
Coherent-prose comparator
A natural follow-up question is whether TE’s advantage holds against a coherent-prose summary
at the same budget rather than against trivial controls. We run the same encoder under a freeprose summary prompt and post-truncate each
summary to TE’s per-row token budget. This
matched-budget contrast isolates representation
format from compression ratio: encoder, consumer, and budget are held fixed; only the surface
form of the compressed passage changes.
4
Experimental Setup
4.1
Datasets and consumer model
We evaluate on three multi-hop QA benchmarks
with distinct depth profiles: MuSiQue (Trivedi
et al., 2022) (n = 2,417 questions; hops ∈
{2, 3, 4}), 2Wiki (Ho et al., 2020) (n = 1,500;
balanced 500 per hop level), and HotpotQA (Yang
et al., 2018) (n = 1,000; all 2-hop). The consumer is Qwen-3.5-9B (base model; HF eager
bf16, greedy decoding) throughout. The encoder
for TE is Claude Sonnet 4.6 via AWS Bedrock
batch inference. Retrieved passages are the goldplus-distractor contexts released with each benchmark; NL, TE, and LLMLingua-2 all read the
same passage set per question.
Auxiliary mechanism: matched-budget
controls
We separately pre-registered an auxiliary mechanism observation (Appendix H): TE performs
semantic-preserving compression only where natural language carries redundancy to strip. The
test is whether TE outperforms three trivialcompression controls at matched per-row token
budget. Each control is computed against TE’s
per-row qwen-token count and rules out a specific
alternative explanation:
4.2
• Character-density. The NL passage is
rescaled at the character level so that its
qwen-token footprint matches TE’s. Rules
out the hypothesis that TE’s gain comes from
per-row character-density manipulation.
Prompt and answer extraction
All representations share a single neutral consumer prompt (Appendix B). The answer is extracted from the consumer’s final-line output by
regex and evaluated against the gold answer with
token-level F1, with binary correctness at F 1 ≥
0.5.
• End-truncation. The NL passage is truncated from the end at the qwen-token boundary so it has TE’s per-row token count. Rules
out the hypothesis that the NL tail is dispensable.
4.3
Baselines
• NL. Full retrieved passages, unmodified. The
standard no-compression baseline.
• Random subsampling. A fixed-seed uniform subsample of qwen-token positions is
drawn from the NL passage, sized to TE’s
budget. Rules out the hypothesis that any
random subset of NL tokens would suffice.
• LLMLingua-2 at rate-50 (Pan et al., 2024).
The primary learned-token-scoring baseline.
• Three matched-budget controls (characterdensity, end-truncation, random subsampling), each sized to TE’s per-row qwentoken count. Defined in §3.3 and Appendix D.
The direction of the controls is pre-registered
(TE should beat all three); the descriptive “9 of
3
Text of page 4
• Coherent-prose summary produced by the same encoder and post-truncated to TE’s perrow token budget (§3.4). 4.4 We use paired-bootstrap 95% confidence intervals over questions (n boot = 10,000, seed 0) for landmark and matched-budget contrasts. The preregistered depth-interaction tests use a binomial GLM with cluster-robust standard errors by question, pooled across datasets by random-effects meta-analysis, with false-discovery-rate correction across the primary interaction family (Benjamini and Hochberg, 1995; DerSimonian and Laird, 1986). Full statistical specifications, the pre-committed heterogeneity rule, and the sensitivity check against a random-intercept fit are in Appendix A. 4.5 We filed a pre-registration prior to data collection. The primary depth-interaction hypothesis, the matched-budget mechanism observation, the FDR correction family, the random-effects pooling specification, the heterogeneity branch rule, and the landmark kill-gate tests are all preregistered; the protocol and amendment chain are reproduced in Appendix H. 4.6 To check that the matched-budget mechanism is not specific to one consumer family, we re-run the three matched-budget controls on a second consumer, Mistral-7B-Instruct-v0.3, on MuSiQue. Long-context NL rows do not fit at 24 GB on this consumer, so the cross-architecture claim rests on the matched-budget controls (where TE and the controls have similar lengths) rather than on the full-NL landmark; details in §6.1. 5 Results 5.1 Matched-budget mechanism: TE beats every trivial control Statistical protocol Control Dataset TE − ctrl (pp) [95% CI] Char-density Char-density Char-density End-truncation End-truncation End-truncation Random-subset Random-subset Random-subset MuSiQue 2Wiki HotpotQA MuSiQue 2Wiki HotpotQA MuSiQue 2Wiki HotpotQA +16.6 [+14.7, +18.4] +13.6 [+11.3, +15.9] +18.2 [+15.2, +21.0] +18.8 [+16.9, +20.6] +15.1 [+12.7, +17.4] +18.6 [+15.7, +21.6] +18.4 [+16.5, +20.3] +16.3 [+13.9, +18.6] +20.2 [+17.3, +23.1] Table 1: Matched-budget mechanism. At matched per-row qwen-token budget, TE outperforms all three trivial controls on all three datasets. Paired bootstrap over questions, n boot = 10,000, seed 0. A6: char-density (density-matched) A7: NL end-trunc (tail dispensable?) A8: random-trunc (random subset?) 25 F1 minus control F1 (pp) Pre-registered protocol 20 +16.6 +18.2 +18.6 +18.8 +18.4 +13.6 15 +15.1 +20.2 +16.3 10 5 0 MuSiQue 2Wiki HotpotQA Figure 1: Matched-budget mechanism. TE beats character-density, end-truncation, and randomsubsampling controls on every dataset at matched perrow token budget. Error bars: 95% paired-bootstrap CI. advantage. Character-density rules out per-row character manipulation. End-truncation rules out the hypothesis that the NL tail is dispensable. Random subsampling rules out the hypothesis that any size-matched random subset of NL tokens would suffice. We read the result as supporting the pre-registered mechanism: TE compresses where natural language carries redundancy, and trivial alternatives that strip surface tokens but do not reexpress content lose the bridge entities a multi-hop reader needs. Cross-architecture replication 5.2 Landmark comparisons at full budget Table 2 reports paired-bootstrap 95% CIs for TE against full-budget NL and against LLMLingua-2 at rate-50. On MuSiQue, the deepest-hop dataset, TE beats NL by +4.75 pp and LLMLingua-2 by +7.67 pp (both p < 0.001). On 2Wiki the TE–NL gap is +1.53 pp with a CI that narrowly spans zero. On HotpotQA the sign flips: TE loses to NL by −2.34 pp (p < 0.01). The between-dataset ordering—positive significant on MuSiQue, positive null on 2Wiki, negative significant on HotpotQA—is the empirical pattern that motivates the post-hoc moderator analysis in Table 1 reports paired-bootstrap 95% confidence intervals for TE versus the three matched-budget controls on each of MuSiQue, 2Wiki, and HotpotQA. All nine intervals exclude zero from above, with point estimates ranging from +13.6 to +20.2 percentage points (Fig. 1). The three controls jointly rule out three specific alternative explanations for TE’s matched-budget 4
Text of page 5
Dataset TE−NL (pp) [95% CI] TE−LLMLingua-2 (pp) [95% CI] MuSiQue +4.75 [+3.06, +6.41] 2Wiki +1.53 [−0.17, +3.24] HotpotQA −2.34 [−4.41, −0.34] +7.67 [+6.00, +9.32] +1.60 [−0.13, +3.36] −2.89 [−4.82, −0.92] Table 2: Landmark paired comparisons. Full-NL and LLMLingua-2 rate-50 budget regimes. Paired bootstrap over questions, n boot = 10,000, seed 0. Per-dataset n: MuSiQue 2,417; 2Wiki 1,500; HotpotQA 1,000. Dataset TE−prose (pp) [95% CI] MuSiQue +11.94 [+10.09, +13.72] 2Wiki +0.32 [−1.69, +2.30] HotpotQA +1.96 [−0.28, +4.37] −4.28 [−6.13, −2.43] +1.28 [−0.77, +3.33] −4.85 [−7.09, −2.69] 5.4 We fit the pre-registered depth-interaction model per dataset, pool MuSiQue and 2Wiki by randomeffects meta-analysis, and apply false-discovery-rate correction across the family. Table 4 reports the four within-dataset slopes plus their metaanalytic pool. None is statistically significant: FDR-adjusted p-values range 0.41 to 0.92 and the pooled meta slopes are indistinguishable from zero (I 2 = 0% on both interaction terms, so the NL (reference) LLMLingua-2:hop_c est [95% CI] p FDR −0.064 [−0.149, +0.021] −0.004 [−0.085, +0.077] −0.033 [−0.091, +0.026] 0.413 0.918 0.413 (Telegraph English) MuSiQue F1 (%) LLMLingua-2 2Wiki 70 50 65 45 60 40 55 35 50 45 30 2 Coherent-prose comparator at matched budget A natural concern with the matched-budget mechanism is whether TE’s advantage persists against a coherent-prose summary at the same budget, rather than against trivial controls that corrupt surface structure. We run Claude Sonnet 4.6 as a coherent-summary encoder on the same three datasets and post-truncate each summary to TE’s per-row qwen- token budget. Table 3 reports the comparison. On MuSiQue, the deepest-hop dataset, TE outperforms the matched-budget coherent prose by +11.94 pp with the CI strictly positive. On 2Wiki and HotpotQA the difference is null. Coherent prose itself loses to full NL on MuSiQue and HotpotQA, suggesting that matched-budget truncation of coherent summaries is a costly operation when the budget is tight: at MuSiQue’s budget, the truncated summary retains roughly half of TE’s named entities, while the untruncated summary at 1.59× the budget retains 78% (Appendix J). 0.711 0.711 0.711 Table 4: Depth-interaction regression. All four withindataset slopes are direction-consistent with the preregistered negative prediction but non-significant after FDR correction; pooled estimates are indistinguishable from zero, I 2 = 0%. Appendix I; we report it there because at n = 3 datasets we cannot rule out unobserved datasetconstruction factors. 5.3 p FDR −0.018 [−0.113, +0.077] −0.019 [−0.113, +0.075] −0.018 [−0.085, +0.048] MuSiQue 2Wiki Meta Table 3: Coherent-prose comparator at matched per-row qwen-token budget. Paired bootstrap over questions, n boot = 10,000, seed 0. Per-dataset n: MuSiQue 2,417; 2Wiki 1,500; HotpotQA 1,000. TE:hop_c est [95% CI] MuSiQue (n=2,417) 2Wiki (n=1,500) Meta (RE) Scope prose−NL (pp) [95% CI] prose−LLMLingua-2 (pp) [95% CI] −6.71 [−8.58, −4.85] +1.21 [−0.80, +3.23] −4.30 [−6.60, −2.12] Scope 3 hop count 4 40 2 3 hop count 4 Figure 2: Depth-interaction null. Per-hop F1 for NL, TE, and LLMLingua-2 within MuSiQue and 2Wiki, with 95% bootstrap CI shading. The TE–NL gap is flat across hop counts, not growing with depth. HotpotQA is excluded because all questions are 2-hop. pre-committed heterogeneity rule did not fire). All four point estimates are negative—directionconsistent with the pre-registered prediction—but at these p-values the direction-consistency is indistinguishable from noise, not weak supporting evidence. Section 6.2 quantifies this with a minimumdetectable- effect-size analysis. Figure 2 visualises the null: the TE–NL gap is flat or non-monotone across hop levels within both MuSiQue and 2Wiki, not rising with depth as the pre-registered mechanism would predict. 5.5 Methodological note Our pilot swept three consumer-prompt variants. Under one variant (role-prompted), TE’s MuSiQue F1 varied by tens of points across seeds relative to the default prompt because verbose responses interact with the normalised-multiset F1 scorer in a way that is orthogonal to context compression. Main-body numbers use the neutral default prompt throughout; the full prompt-by-metric table is in Appendix F. Depth interaction is null 6 Discussion 6.1 What the matched-budget result shows The matched-budget mechanism result is the central finding of this paper. At the same per-row 5
Text of page 6
able. The depth-dependent effect may exist but our sample is too small to detect it. Or the depth-dependent mechanism may simply be absent: TE’s advantage may be a constant offset over compressed-NL equivalents rather than a depthdependent widening over full NL. A minimum-detectable-effect-size analysis bounds the interpretation. The pooled meta slope is −0.018 log-odds per hop with cluster-robust standard error 0.034; at 80% power and α = 0.05 two-sided, the minimum detectable slope is roughly 0.095 log-odds per hop. Translated to F1 at our consumer’s MuSiQue baseline, that corresponds to a TE–NL widening of about 4–5 percentage points across the hop-2-to-hop-4 range. The design therefore rules out widenings of that magnitude with ≈ 80% power but cannot distinguish smaller effects from no effect. The null is informative against a strong-form depth-dependent advantage and uninformative about a weak-form one. What our data do support is that the TE–NL gap does not grow monotonically with hop count within MuSiQue or 2Wiki at any appreciable magnitude (Fig. 2). qwen-token budget, TE outperforms characterdensity, end-truncation, and random-subsampling controls by +13.6 to +20.2 percentage points across three datasets. The three controls exhaust three mechanistically distinct trivial-compression strategies—compress by character deletion, drop the NL tail, drop random NL tokens—and none preserves the bridge entities a multi-hop reader needs. TE preserves bridge entities verbatim and replaces connective NL tissue with symbolic operators, a re-expression of the same semantic content at the same budget. We read this as evidence that TE compresses where natural language carries redundancy: the matched-budget gain is the benefit of that operation rather than an artefact of the compression strategies the controls implement. The coherent-prose comparator sharpens the reading. Coherent prose at matched budget retains roughly half of TE’s named entities on MuSiQue, while a longer untruncated coherent summary at 1.59× the budget retains 78%. The lost entities are bridge entities the matched-budget truncation drops, and that loss is a property of coherent prose at tight budgets, not an artefact of the encoder. TE holds entity content more densely than either NL or coherent prose because pipe-separated triples reserve every token for entity-bearing content that articles, auxiliaries, and connectives would otherwise consume. We name this the density argument: at matched budget, the relevant axis is entities-per-token, and TE’s representation is the denser one. The matched-budget mechanism replicates on a second consumer family. Re-running the three controls on Mistral-7B-Instruct-v0.3 on MuSiQue yields paired-bootstrap 95% CIs strictly positive on all three controls (+8.62/ + 11.08/ + 10.62 pp), direction-consistent with Qwen-3.5-9B at the smaller magnitude expected for a 7B consumer. The full-NL landmark is not a valid baseline on Mistral because long-context NL rows do not fit at 24 GB and the surviving subset is biased toward shorter (easier) questions; the cross-architecture claim therefore rests on the matched-budget controls only. 6.2 The depth-interaction hypothesis predicted that TE’s advantage over NL would grow with hop count within multi-hop datasets. All four withindataset slopes came out direction-consistent but non-significant, and the pooled estimates are indistinguishable from zero. Two readings are avail- We tested Telegraph English, a readable symbolic re-expression of retrieved passages, as a matchedbudget alternative to natural language for multihop question answering with small language models. Across three benchmarks, TE outperforms three trivial-compression controls and a coherentprose summariser at the same per-row token bud- 6.3 Implications and encoder provenance Our experiments use Claude Sonnet 4.6 as the TE encoder, deliberately chosen as a strong frontier model so that the consumer-reading question is not confounded with translator capacity. A natural follow-up is whether a much smaller encoder, given domain-matched training, could substitute. A pilot we describe in Appendix L fine-tunes Qwen-3.5-0.8B on 5,000 entity-preserving rows with oracle bridge- entity conditioning, evaluates on the held-out MuSiQue test shard, and beats both LLMLingua-2 and the Sonnet TE encoder at roughly 10% of the qwen-token budget. The pilot is single-dataset and uses oracle conditioning, so it is an upper bound on encoder substitutability rather than a deployable alternative; the general translator-capacity question remains open. 7 Why the null is informative 6 Conclusion
Text of page 7
get, with all nine matched-budget confidence intervals strictly positive. A pre- registered depthinteraction hypothesis is null: the advantage does not grow with reasoning depth, and a minimumdetectable-effect-size analysis bounds the design to ruling out widenings of 4–5 F1 points across hop levels. We interpret these results as evidence that the operative property of TE is entity density per token, and the natural follow-up is whether smaller encoders, evaluated without oracle entity conditioning, can preserve that density. 8 Limitations and Broader Impact 8.1 Limitations Between-dataset gradient is n = 3. The posthoc reading of the between-dataset TE–NL gradient as correlating with dataset NL ceiling rests on three datasets. NL-ceiling- proximity is the only monotone candidate moderator at this n; unobserved dataset-construction factors (bridge-entity retrievability, distractor-passage content) correlated with NL ceiling cannot be ruled out. A confirmatory replication would require a fourth dataset where NL ceiling varies independently of dataset construction, and we report the gradient as an exploratory observation only (Appendix I). encoders. Coherent prose loses entity coverage at matched budget. The matched-budget coherentprose comparator truncates the encoder’s freeprose output to TE’s per-row budget. On a 10-row MuSiQue sample this drops entity coverage from 0.782 (raw, 1.59× budget) to 0.497 (matched budget). That is a density property of coherent prose at tight budgets, not an engineering artefact we can remove. A reader whose deployment budget is measured in retained entities rather than consumer tokens should treat our prose-comparator evidence as upper-bounded; a comparison at matched entity coverage would allow the prose representation a larger budget and would answer a different question. TE-at-larger-budget counterfactual untested. We compare TE at matched qwen-token budget to coherent prose at the same budget; we do not test TE at 1.59× budget against coherent prose at 1.59× budget. The latter would isolate representation- format advantage from compression-ratio advantage and remains followup work. Prompt-template control. A control on MuSiQue (sub-sample n = 500) under the role-prompted consumer prompt confirms the methodological-artefact reading of §5.5: NL F1 (7.61%) and TE F1 (7.14%) collapse together under that prompt, so the collapse is prompt-template-universal and not specific to TE. Under the explicit-reasoning prompt on the same sub-sample, TE F1 (34.28%) substantially exceeds NL F1 (3.12%), suggesting TE is more prompt- robust than raw NL under that variant. We flag this as suggestive only because the sub-sample is MuSiQue at n = 500. Cross-architecture coverage is mechanism-only. The matched-budget mechanism is tested on two consumer families (Qwen-3.5-9B and Mistral-7B- Instruct-v0.3) and is direction-consistent on both (§6.1). The aggregate TE–NL landmark and the dataset-hardness gradient, however, are tested only on Qwen-3.5-9B; the Mistral full-NL baseline is biased by long-context out-of-memory errors on the NL rows, so we do not claim architecture robustness for the landmark or for the betweendataset gradient. The coherent-prose comparator was not run on Mistral for the same reason. Crossfamily landmark coverage at a consumer with sufficient context-window headroom is follow-up work. Metric brittleness. We use token-level F1 with the standard normalised-multiset implementation. F1 is sensitive to response verbosity (the roleprompt artefact in §5.5 is a manifestation). We report exact-match (EM) alongside F1 on MuSiQue in Appendix G; EM and F1 deltas track the same direction across all five comparators, which lends confidence that the depth- interaction null and the matched-budget mechanism are not F1-scorer artefacts. One divergence point is worth flagging: TE’s aggregate EM advantage over NL on MuSiQue (+1.08 pp, 95% CI [−0.58, +2.69]) is substantially smaller than its F1 advantage Single encoder model and frozen prompt. TE’s encoder (Claude Sonnet 4.6 via AWS Bedrock batch inference) and the encoder prompt are fixed. We do not vary encoder capacity, encoder family, or prompt phrasing. A dedicated analysis of encoder sensitivity is future work; the present paper is about whether a frozen tokenizeraware encoder can serve as a matched-budget alternative to NL, not about the landscape of such 7
Text of page 8
(+5.24 pp, 95% CI [+3.59, +6.89]) and the EM CI crosses zero, consistent with the aggregate TE– NL gap being partly carried by partial-credit tokens that F1 counts as recall and EM counts as a miss. A full LLM-judge sensitivity sweep remains future work. 8.2 TE is a context-compression representation that, in our data, preserves bridge-entity spans from NL verbatim; it is therefore no more vulnerable to entity leakage than retrieval over the source NL passages themselves. The encoder is frozen and task-agnostic, so TE does not implicitly encode downstream task supervision in its output—a property we consider favourable for auditability. We see no application-specific harms unique to TE relative to the NL baseline it is meant to replace. Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics. The depth-interaction null is bounded, not erased. The minimum-detectable-effect-size analysis in §6.2 shows our pooled design was powered for per-hop interaction slopes ≳ 0.095 log-odds/hop and underpowered for slopes below that. Readers should treat the null as informative against strong-form depth- dependent widening (slopes ≳ 4–5 pp across the hop-2-to-hop-4 range) and uninformative about weak-form widening. Closing the gap would require either substantially larger n per dataset or a narrower prediction. Passage-set scope. We use the gold-plus-distractor contexts released with each benchmark as retrieval input. This isolates the representationvs-NL question from the retrieval question, but it also means our results do not speak directly to deployed retrieval pipelines where distractor quality varies. Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: Program-aided language models. In Proceedings of the 40th International Conference on Machine Learning. Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2024. LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Jesse Mu, Xiang Lisa Li, and Noah D. Goodman. 2023. Learning to compress prompts with gist tokens. In Advances in Neural Information Processing Systems 36. Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114. Broader impact References Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1):289–300. Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Rebecca DerSimonian and Nan Laird. 1986. Metaanalysis in clinical trials. Controlled Clinical Trials, 7(3):177–188. Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. 2024. LLMLingua- 2: Data distillation for efficient and faithful taskagnostic prompt compression. In Findings of the Association for Computational Linguistics: ACL 2024. Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35. Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024. RE- COMP: Improving retrieval-augmented LMs with context compression and selective augmentation. In Proceedings of the Twelfth International Conference on Learning Representations. 8
Text of page 9
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and
Christopher D. Manning. 2018. HotpotQA: A
dataset for diverse, explainable multi-hop question
answering. In Proceedings of the 2018 Conference
on Empirical Methods in Natural Language Processing.
Pre-committed heterogeneity rule. Before fitting, we committed to a heterogeneity branch rule:
if I 2 > 75% on a pooled slope, pooling is not
interpretable and we fall back to per-dataset inference. The rule did not fire in our sample (I 2 = 0%
on both pooled interaction slopes).
Howard Yen, Tianyu Gao, and Danqi Chen. 2024.
Long-context language modeling with parallel context encoding. In Proceedings of the 62nd Annual
Meeting of the Association for Computational Linguistics.
Chanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong, and Jaewoo Kang. 2024. CompAct:
Compressing retrieved documents actively for question answering. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language
Processing.
Paired-bootstrap. For landmark and matchedbudget contrasts, paired-bootstrap 95% confidence
intervals resample question ids with n boot =
10,000 and seed 0, using the percentile bracket
on the bootstrap distribution.
A
The pre-registered primary hypothesis predicts
that TE’s advantage over NL grows with
reasoning depth, operationalised as a negative representation × hop_count interaction within multi-hop datasets. Let y i ∈
{0, 1} be the correctness of the consumer’s answer to question i under representation r i ∈
{NL, TE, LLMLingua-2} at hop count h i ∈
{2, 3, 4}, with y i = 1 iff token-level F1 against
the gold reference is ≥ 0.5. We fit, per dataset, a
binomial GLM with logit link
B
Consumer prompt
The default consumer prompt used for all mainbody numbers is reproduced verbatim below, with
{context} and {question} placeholders substituted at runtime:
Statistical protocol
C
Answer the question using only
the context. Return the answer
on the final line.
Context:
{context}
Question: {question}
Answer:
Telegraph English encoder prompt
and example
TE is produced by Claude Sonnet 4.6 (via managed batch inference) with a fixed prompt that
instructs entity-preserving re-expression at a target qwen-token budget. The encoder prompt is
reproduced in full in the supplementary materials;
an illustrative pre/post pair from a MuSiQue row
is:
logit P(y i = 1) = α+β r r i +β h hop c +γ r,h (r i ·hop c )
(1)
where hop c = h i − h̄ centers hop count at its
dataset mean and cluster-robust standard errors by
question_id serve as a practical proxy for the
pre-registered glmer + (1|question_id)
random-intercept specification (the two specifications agree on all inferential conclusions; sensitivity analysis in Appendix E). The pre-registered
prediction is γ TE,h < 0. HotpotQA is excluded
because all its questions are 2-hop and the withindataset slope is undefined.
NL: “Barack Obama was born in
Honolulu, Hawaii. Honolulu
is the capital of the state of
Hawaii. Hawaii is a state in
the United States.”
TE: “Barack Obama @born
Honolulu | Honolulu @capital_of
Hawaii | Hawaii @state_in
United States.”
TE preserves the entity spans (“Barack
Obama”, “Honolulu”, “Hawaii”, “United States”)
verbatim and rewrites connective NL tissue into
pipe-separated symbolic clauses with @-prefixed
operators.
Meta-analysis and multiple-comparison correction. Per-dataset interaction slopes are pooled
across MuSiQue and 2Wiki by random-effects
meta-analysis (DerSimonian–Laird; DerSimonian
and Laird, 1986), yielding a pooled point estimate, 95% Wald CI, and I 2 heterogeneity statistic.
False-discovery-rate correction (Benjamini and
Hochberg, 1995) is applied across three tests per
interaction term ({M uSiQue, 2W iki, meta}).
D
A6/A7/A8 specification
All three trivial-compression controls are operationalised per row against TE’s per-row qwentoken count. Denote by B i the qwen-token count
of TE’s passage for question i.
9
Text of page 10
A6 (char-density). We rescale the NL passage by a character-level density transform: drop every k-th character with k chosen per row so that the post-transform qwen-token count matches B i . Whitespace is preserved to keep the output humanreadable. E F Wave-1a pilot numbers on MuSiQue under three consumer prompts (default, explicit_reasoning, role_prompted) are tabulated in Table 5 below. Under role_prompted, TE’s F1 varies by up to 18.3 pp relative to default on the same rows, driven by response-length interaction with the normalised-multiset F1 scorer. G Token-level F1 with the standard normalisedmultiset implementation is sensitive to response 42.74 41.10 24.43 37.99 36.47 34.80 ∆ +4.75 +4.63 −10.37 Table 5: Prompt-format sensitivity on MuSiQue (Wave- 1a pilot). The role_prompted row shows a large negative swing for TE driven by verbose-response F1 deflation, not by an intrinsic compression loss. Mainbody numbers throughout the paper use default. verbosity (§F). To check that the C1 ′ null and the C2 confirmation are not F1-scorer artefacts, we re-score the same MuSiQue predictions with exactmatch (EM) and report paired-bootstrap 95% CIs on each representation’s EM delta versus NL at matched budget. The MuSiQue evaluation shard contains both F1 and EM for every row; no new inference is required. Cluster-robust GLM vs. glmer sensitivity Our primary C1 ′ specification uses a binomial GLM with cluster-robust standard errors by question_id as a practical proxy for the pre-registered glmer(correct ~ representation * hop_c + random-intercept (1|question_id)) specification. We re-fit the full glmer specification on both MuSiQue and 2Wiki as a sensitivity analysis; the estimated representation × hop_c slopes agree with the cluster-robust GLM to three decimal places on MuSiQue and two decimal places on 2Wiki, and the inferential conclusion (all four slopes direction-consistent, none significant after FDR-BH) is unchanged. default explicit_reasoning role_prompted A7 (NL end-truncation). We truncate the NL passage from its end at the qwen-token boundary so that the truncated passage has B i qwen tokens. Partial final words are dropped at the nearest word boundary to avoid UTF-8 fragments. A8 (NL random-subset). We sample B i qwentoken positions uniformly without replacement from the NL passage (fixed seed 42) and concatenate the selected tokens with a single space between segments. All three operate in qwen-token space, not word or character space, so the budget matches what the consumer LM actually sees at its tokenizer. TE F1 (%) NL F1 (%) Prompt Cond NL TE LLMLingua-2 A6 A7 A8 F1 mean EM mean (%) (%) 37.50 42.74 35.07 25.38 23.96 24.31 F1 ∆ vs NL (pp, 95% CI) EM ∆ vs NL (pp, 95% CI) 28.09 — — 29.17 +5.24 [+3.59, +6.89] +1.08 [−0.58, +2.69] 25.78 −2.43 [−3.89, −1.00] −2.32 [−3.68, −0.99] 17.29 −12.12 [−14.03, −10.20] −10.80 [−12.66, −8.94] 16.34 −13.54 [−15.41, −11.66] −11.75 [−13.57, −9.93] 17.34 −13.19 [−15.09, −11.30] −10.76 [−12.58, −8.94] Table 6: F1 and EM on MuSiQue (n=2,417 questions paired) with paired-bootstrap 95% CIs on the delta vs NL (n boot = 10,000, seed 0). EM and F1 deltas are direction-consistent across all five comparators. Note that TE’s EM advantage over NL is small (+1.08 pp) and its 95% CI crosses zero, whereas its F1 advantage (+5.24 pp) is clearly positive. The three trivialcompression controls (A6/A7/A8) show large negative deltas on both metrics with CIs well below zero, and LLMLingua-2 is negative on both metrics with CIs excluding zero. The direction-consistency across F1 and EM lends confidence that the C2 mechanism finding (A6/A7/A8 all far below TE at matched budget) is not an F1-scorer artefact. On the positive TE >NL comparison, the EM CI that crosses zero is informative: it suggests TE’s aggregate advantage on MuSiQue is at least partly carried by partialcredit rewards where TE emits the correct bridge entity alongside additional tokens that F1 counts as recall but EM counts as a miss. This is consistent with the compression-where-redundancy interpretation in §6.1 (which is about mechanism at matched budget, not about the size of the aggregate F1 advantage). A full LLM-judge sensitivity sweep remains future work (§8.1). Prompt × F1-metric artefact EM alongside F1 on MuSiQue 10
Text of page 11
H
Pre-registration protocol and
amendments
These pre-commitments—FDR-BH correction
across {MuSiQue, 2Wiki, meta}, a heterogeneity rule (I 2 > 75%; it did not fire), an analytic
MDES bound reported alongside the null, and
explicit hypothesis-generating-only labelling of
the between-dataset moderator—jointly prevent
auxiliary-to-primary promotion after the null primary.
The pre-registration was filed prior to data collection (DOI and filing date omitted from the submission version for reviewer anonymization; restored in the camera-ready). The primary C1 ′ hypothesis, the auxiliary C2 mechanism observation
(§11.7 obs 2), the FDR-BH correction family, the
DerSimonian–Laird pooling specification, the precommitted I 2 > 75% heterogeneity rule, and the
landmark kill-gate tests K-F1-1, K-TA-1, and K-γ-
1 are all pre-registered. Two in-repo amendments
were filed during the study:
A full amendment log with SHA-stamped timestamps is included in the supplementary materials.
I
“On HotpotQA (predominantly 1–2 hop), TE
does not beat NL: ∆F 1(NL − TE) > 0. This
validates that the interaction is emergent with
depth, not a global TE-wins effect.” (preregistered document, §4)
MuSiQue
4
2
2Wiki
0
2
HotpotQA
n=3 datasets
post-hoc moderator
hypothesis-generating only
4
6
30
40
50
60
70
NL F1 (%) [dataset NL-ceiling proxy]
80
Figure 3: Post-hoc NL-ceiling-proximity observation (n = 3 datasets, hypothesis-generating only).
Between-dataset TE−NL ∆F1 (pp) plotted against
dataset NL F1 (%). With only three datasets any monotonic moderator fits similarly; we therefore show no
fitted line, no R 2 , and no coefficient. The ordering is
consistent with—but does not confirm—an NL-ceiling-proximity reading; unobserved dataset-construction
factors correlated with NL ceiling cannot be ruled out
at n = 3. Error bars: 95% paired-bootstrap CI on ∆F1.
Post-hoc observation: between-dataset
gradient
This appendix expands the brief main-body
pointer at the end of §5.4 into the full detail of
the between-dataset TE−NL gradient. We report
the gradient as an exploratory observation only,
not as a contribution.
The between-dataset ordering in Table 2
(MuSiQue +4.75 significant, 2Wiki +1.53 null,
HotpotQA −2.34 significant negative) is directionconsistent with the pre-registered P2b verbatim:
6
• §11 amendment: filed C2 (compressionwhere-redundancy, §11.7 obs 2) as a separately pre-registered auxiliary mechanism,
prior to the Wave-2 cloud run that produced
the A6/A7/A8 data reported in Table 1.
8
• §10 amendment: documented a defaultprompt bug discovered in the Wave-1a pilot
and specified the remedial Wave-1b rerun at
matched LLMLingua-2 budget.
The direction is pre-registered; the specific moderator that explains the gradient is not. In our
n = 3 datasets the ordering correlates monotonically with dataset NL F1 (MuSiQue 38.0,
2Wiki 56.0, HotpotQA 69.7). Hop-based moderators (modal hop count, mean hop depth, 4-hop
share) do not have the monotone shape required:
modal hop counts are 2/no-mode/2 respectively,
mean hop depths are 2.65/3.0/2.0, and 4-hop
shares are 17%/33%/0%. NL-ceiling-proximity
is therefore the only monotone candidate moderator at n = 3. We report it as a post-hoc
observational pattern only—unobserved datasetconstruction factors (e.g., bridge-entity retrievability, distractor-passage content) correlated with NL
ceiling cannot be ruled out. Figure 3 shows the
three points with no fitted line and an explicit
n = 3 caveat.
NL F1 (pp)
At n = 3 datasets, NL-ceiling-proximity is the
only monotone candidate moderator; unobserved
dataset-construction factors (e.g., bridge-entity retrievability, distractor-passage content) cannot be
ruled out. We therefore do not treat the gradient
as a finding of this paper. A confirmatory replication with a fourth dataset where NL ceiling varies
independently of dataset construction is the minimum prerequisite for any claim; we reserve that
for follow-up work.
Hop distribution per dataset (reported for completeness, not a candidate moderator at n =
11
Text of page 12
3). MuSiQue: 2-hop 1,252 (51.8%), 3-hop 760 (31.4%), 4-hop 405 (16.8%); modal 2, mean 2.65. 2Wiki: 2-hop 500, 3-hop 500, 4-hop 500 (balanced); no single mode, mean 3.0. HotpotQA: 2-hop 1,000 (100%); modal 2, mean 2.0. None of modal hop, mean hop, or 4-hop share is monotone with the between-dataset TE − NL gradient. J over NL on any dataset (∆ = −6.6, −13.4, −15.7 pp on MuSiQue, 2Wiki, HotpotQA respectively); F2-struct is cut per pre-registration. K-γ-1, a hop- 3 rescue frame intended to outperform TE on 3hop questions in a revise-and-review pass, fails: combined MuSiQue + 2Wiki hop-3 ∆(γ − TE) = −0.15 pp, 95% CI [−0.78, +0.48], n = 1,260. γ is cut. Both negatives are useful: F2-struct falsifies a “surface structure alone suffices” hypothesis, and K-γ-1 falsifies a specific “deeper revision pass rescues the hard hops” hypothesis. F3 coherent-prose comparator: entity coverage and landmark contrasts F3 evaluates Claude Sonnet 4.6 as a frontier coherent-summary encoder on the same three datasets, post-truncated to each per-row TE qwentoken budget (A7-style), so F3 is evaluated at the same budget as TE on identical consumer rows. Entity-coverage QC on a 10-row random MuSiQue sample (seed 0): at matched TE qwentoken budget F3 retains 0.497 of TE’s named entities, versus 0.782 for F3 raw (pre-truncation) at 1.59× TE’s budget and 0.824 for LLMLingua-2 rate-50 at 2.29× TE’s character budget (Table 7); the drop 0.782 → 0.497 is produced entirely by matched-budget truncation, not by the encoder. Comparing F3 vs A7 on MuSiQue (both truncated to TE’s per-row qwen-token budget, differing only in summarization quality) gives F3−A7 ≈ +7 pp: coherent summarization adds value over raw truncation even after both lose entity coverage, so the TE−F3 contrast in Table 3 isolates the density advantage from the summarization-quality advantage. Paired-bootstrap 95% CIs: n boot = 10,000, seed 0, paired by question_id; perdataset n: MuSiQue 2,417; 2Wiki 1,500; HotpotQA 1,000. Compressor TE (Telegraph English) F3 (coherent prose, truncated) F3 raw (pre-truncation) LLMLingua-2 rate-50 Entity coverage (median) Budget regime 1.000 † 0.497 0.782 0.824 matched TE qwen-tokens matched TE qwen-tokens 1.59× TE qwen-tokens 2.29× TE char budget L Follow-up pilot: small-LM encoder on MuSiQue with oracle entity conditioning Motivation. §6.3 chose Claude Sonnet 4.6 as the TE encoder to isolate the consumer-reading question from translator-capacity confounds. A reviewer may reasonably ask whether a much smaller encoder, given domain-matched entitypreserving training, could substitute. This pilot is a single-dataset proof-of-concept for encoder substitutability; it is not a full replication of the main-experiment matrix and does not close the general translator-capacity question. Setup. We fine-tuned Qwen-3.5-0.8B (base) via LoRA on 5,000 entity-preserving Wikipedia QA rows drawn from the MuSiQue and HotpotQA training splits (2,500 each). Each training row pairs a question, a retrieval passage, the answer, and the decomposition’s bridge entities (the set of surface forms the decomposition labels as appearing in intermediate hops). The encoder is conditioned at both train and inference on the gold bridge-entity set as a system-prompt side channel; the supervised target is a Telegraph-English paraphrase of the passage produced by Sonnet 4.6 under a bridge-entity-aware prompt. At evaluation we re-translate the held-out MuSiQue n=2,417 test shard through the fine-tuned encoder under the same oracle conditioning, then send the output to the main Qwen-3.5-9B consumer using the same default prompt, decoder setting, and F1 scorer as the main experiments. Table 7: Entity-coverage QC on a 10-row random MuSiQue sample (seed 0). † TE is the reference by construction. F3 at matched qwen-token budget retains about half of TE’s named entities; the ∼30-pp gap from F3 raw is produced by matched-budget truncation, not by the Sonnet encoder. LLMLingua-2’s higher coverage is measured at a different (more permissive) budget regime. K Results. Evaluating all four conditions on the same n=2,417 held-out items and pairing on question id: Honest negatives: F2-struct and γ-rescue Two pre-registered kill gates fired. F2-struct (a structured NL reformulation) does not improve 12 • E10-v2 (retrained 0.8B): F1 48.04% • TE (Sonnet 4.6, same-job): F1 42.40%
Text of page 13
• NL baseline: F1 37.50% • LLMLingua-2 (rate-50, matched to TE budget): F1 35.07% Paired bootstrap over questions (n boot = 10,000, seed 0): • E10-v2 − LLMLingua-2: +12.97 pp, 95% CI [+11.00, +14.92] • E10-v2 − TE: +5.64 pp, [+3.76, +7.54] • E10-v2 − NL: +10.54 pp, [+8.55, +12.53] All three CIs are strictly positive. The samejob TE F1 42.40% differs by 0.3 pp from the main-experiment TE 42.74%, consistent with deterministic-decode batch-order shifts across consumer-eval runs. binding constraint on our mechanism and density claims: a domain-matched 0.8B encoder with oracle entity conditioning reaches consumer F1 above LLMLingua-2 and above the frontier TE encoder at a fraction of the token budget. It does not establish that translator capacity is unimportant across 2Wiki or HotpotQA, that small-LM encoders match TE without oracle conditioning, or that the density property transfers to matchedbudget generation. The general translator-capacity question remains open and is the subject of concurrent work. Artifacts. The following pilot artifacts accompany this paper as supplementary material: the 5,000-row entity-preserving training data; the Sonnet-generated Telegraph-English targets; the LoRA fine-tune config and training-loss curves; the re-translated MuSiQue n=2,417 test shard produced by the fine-tuned encoder; the consumereval shard, per-row F1/EM outputs, and pairedbootstrap JSON with all three contrasts. Scope limits. Three limitations bound the reading of these numbers. (i) Single-dataset. The pilot tests only MuSiQue; the main paper’s claims span three datasets. Running the pilot on 2Wiki and HotpotQA is mechanical but was deprioritised as the MuSiQue result already suffices to address the “frontier-in, frontier-out” framing concern on the dataset where TE most clearly dominates. (ii) Budget mismatch. The retrained 0.8B output sits at ∼ 10% of TE’s qwen-token budget (median 126 vs. 1251 per row; only 9/2417 rows exceed TE’s budget and are truncated). The +5.64 pp lift over TE therefore reads as an extreme-density data point, not matched-budget dominance. A matched-budget probe (letting E10-v2 generate up to TE’s per-row budget) would disambiguate extreme-density is sufficient from E10-v2 happens to work at low budget; we do not run that probe here. (iii) Oracle conditioning at inference. The decomposition’s bridge-entity set is used as prompting input to the encoder at inference time. In a deployment setting these labels are not available without either human annotation or a separate entity-extraction model; the pilot thus demonstrates an upper bound on encoder substitutability, not a deployable alternative. A non-oracle control (conditioning on NER-predicted bridge entities, or on entities recovered from the consumer’s firstpass guess) is the natural follow-up. What the pilot does and does not establish. Subject to limits (i)–(iii), the pilot establishes that on MuSiQue translator capacity is not the 13