Evaluating Relational Context Compression at Realized Token Budgets
Back to the paper page. ICLR 2027 submission, September 2026.
All 19 pages are shown below.
Text of page 1
E VALUATING R ELATIONAL C ONTEXT C OMPRESSION AT R EALIZED T OKEN B UDGETS Anonymous authors Paper under double-blind review A BSTRACT Context compression must preserve answerable information within a token budget, yet shorter text alone cannot isolate the effect of its representation. We study Telegraph English, which rewrites retrieved passages into relational statements, using a pre-registered encoder–consumer matrix and an open compiler trained by distillation. Six of seven encoders violated the registered output schema on most or all inputs. Only 33 of 9,600 budgeted outputs fell within the registered band: at or below the requested budget and no more than two percent short. These findings motivate evaluating answer F1 at realized token ratios. The matrix was collected, but all registered confirmatory comparisons were unavailable because their comparator or calibration conditions were missing or ineligible. The remaining answer-F1 comparisons with full prose are descriptive and score construction failures as empty answers. The second compiler failed its pre-registered acceptance test against a budget-targeted pruning baseline. Its out-of-distribution component was unscorable because answer keys were absent, and citation-reach judging was not reached. Context-compression comparisons require measured token budgets, explicit construction failures, and available comparator evidence. 1 I NTRODUCTION Context compression changes both the amount of evidence a language model receives and the form in which it receives it. Deletion retains selected parts of a passage; rewriting can express the same relationships in a different syntax. The text actually delivered determines both the reader’s evidence and its token cost. Telegraph English (TE) rewrites source passages into compact relational statements (Arbuzov et al., 2026; Bei et al., 2026). For illustration, “Orin’s adviser is Sana, and Sana works at North Lab” can become “Orin | adviser | Sana; Sana | workplace | North Lab.” The intermediate name connects the relations needed to answer where Orin’s adviser works. An encoder constructs the relations, and a separate consumer reads them to answer the question. Token budgets introduce a second dependency into this comparison. An instruction to produce a short text specifies what the encoder should deliver; tokenization measures what it actually delivered. Two encoders given the same requested budget can therefore present different amounts of context to a reader. A study of relational form needs to follow the produced text through construction, measurement, and answering. Only 33 of 9,600 outputs in our budgeted generation study fall within the registered 98–100% target band. The two anchors’ median lengths exceed their requests throughout the measured grid. A preregistered encoder–consumer matrix places reading outcomes at the delivered lengths and exposes different encoder orderings across readers. Its intended matched rewrite comparison is unavailable, so these observations describe the constructed inputs relative to full prose. A separate distilled compiler supplies the complementary case. Its fixed serializer enforces a token ceiling, but mechanically valid plans still yield low downstream answer scores: the requested-budget area is 0.0658 for the compiler against 0.3936 for the pruning baseline. Together, the studies follow compression from a requested allowance, through the text supplied, to the reader’s answer. This sequence distinguishes requesting a short input, delivering one, and preserving useful evidence within it. 1 Reviewers: please read the Reviewer Guidelines (iclr.cc/Conferences/2027/ReviewerGuidelines) and the AI Policy for Reviewers (iclr.cc/Conferences/2027/AIPolicyForReviewers). If you used AI to expand, edit, or polish your review, please provide the input text to the LLM. Better still, consider skipping the LLM and submitting your original text: we, and the authors, are much more interested in your unedited thoughts than in what an LLM has to say. AI-assisted or not, you are putting your name and reputation behind your review: LLM-generated falsehoods, hallucinations or misrepresentations are subject to disciplinary action, which may include desk-rejecting all papers you have authored.
Text of page 2
2 One source, several ways to supply it to a reader. An encoder rewrites a supplied context, and a consumer answers a question using the resulting text. The matrix contains 1,200 fixed examples from MuSiQue, 2WikiMultiHopQA, and HotpotQA (Trivedi et al., 2022; Ho et al., 2020; Yang et al., 2018). Full prose is that original context, providing a common reference for the information available before rewriting. Changing the representation changes the evidence presented to the reader while preserving the question being asked. An encoder can omit a useful fact, express it in a different form, or fail to supply text; a consumer can also fail to use information that remains. Answer accuracy measures the resulting procedure on the fixed question. The representation and its construction status identify which input produced that answer, allowing the later analysis to relate the score to delivered length and text availability. E VALUATION DESIGN Both studies use a modified token-overlap F1, denoted F1 below. For normalized prediction and answer token lists P, G, let u = | set(P ) ∩ set(G)|. For nonempty lists the score is F1(P, G) = 2u . |P | + |G| The numerator counts distinct shared tokens; the denominators retain repetitions, so this differs from multiset token F1. The scorer takes the highest score over the supplied answer aliases. Appendix A specifies normalization and empty-answer handling. Encoders and consumers. The cross-family matrix is the pre-registered grid of fixed encoder and consumer models evaluated on these examples. Its five encoder families span four provider lineages: GPT, Claude Sonnet, Llama, Magistral, and Ministral; the last two share the Mistral lineage. The encoder instances are GPT-5.6-sol, Sonnet 5, Llama 4 Maverick, Magistral Small, and Ministral 8B. The consumers are Qwen3.5-9B, Llama3.1-8B-Instruct, and Mistral-7B-Instruct-v0.3. These models are fixed design points. Crossing encoders with consumers shows how the same construction procedure is read by different model families. Appendix B records model and construction details. Requested and realized lengths. The realized ratio is the token count of the supplied representation divided by that of full prose, using the same consumer tokenizer. This context-only measure excludes the question and answer instructions. Natural-length Telegraph English receives no requested compression ratio. Budgeted Telegraph English receives a target between 0.25 and 0.70 of the full-prose context, followed by at most one token-count feedback revision. The matrix prompt is question-agnostic, task-informed, and tokenizer-neutral; the budget wrapper supplies the numerical target. Question-agnostic construction sees the source without the particular evaluation question. It must therefore select a representation for subsequent reading without using that question to choose which facts to keep. The prompt requests flat, readable relations and preservation of factual content within a fixed output envelope. Readable relations define the requested surface form; the consumer evaluation measures answer quality from the produced text. Section 3 traces budget requests through to measured outputs. What a representation comparison controls. Full prose and an encoder-produced prose rewrite serve different roles. Full prose anchors answerability to the original context, while a prose rewrite can control for passing the source through an encoder and imposing a length constraint. The registered direct-prose condition uses the same encoder and source as TE and targets its per-row consumertoken budget. It would hold the source, encoder, reader, and delivered token budget constant while comparing two rewriting procedures; their retained facts could still differ. This condition did not produce an eligible comparison, so the available matrix describes reader accuracy for the delivered TE inputs relative to full prose. Failure and comparison rules. The cross-family matrix retains construction failures as empty answers with F1 zero. This measures the complete construction-and-reading procedure, including its inability to supply usable text. The natural-length table pools examples within each encoder– consumer pair; the descriptive paired contrast instead gives equal weight to each dataset, encoder, 2
Text of page 3
and consumer combination. The shared full-prose observation is joined to every relevant comparison by question identifier, retaining pairing through the statistical analysis. A separate compiler study. The open compiler tests a different route to constructing TE: a trained model emits a source-linked plan, and a fixed serializer supplies the text for reading. It has its own training data, construction rules, and acceptance test. It additionally includes QASPER research-paper questions (Dasigi et al., 2021). Its frozen score integrates F1 over the requested budget grid, whereas the matrix frontier uses measured ratios with requested positions assigned to construction failures. The compiler result therefore evaluates the learned construction procedure as it responds to budget requests. Its acceptance decision uses the compiler’s own paired observations; the matrix supplies a separate study of encoder-produced representations. 3 F ROM BUDGET REQUESTS TO DELIVERED CONTEXT Recovering text and satisfying the schema are separate checks. The encoder interface requests an envelope containing a relational-text field and an issues list. Six of seven encoders violate this schema on most or all inputs in the broader generation census. This census includes the matrix encoders and the smaller and larger Ministral variants. GPT-5.6-sol is the sole encoder that supplies every requested envelope correctly. The difficulty is therefore visible before a consumer answers a question. A mechanical recovery pass extracts readable bodies from some malformed envelopes. This produces two useful counts: outputs that obey the interface and outputs that supply text for a reader. For example, Llama’s raw envelopes all violate the schema, while recovery supplies nonempty text for all its inputs (Appendix C). An available body can still omit or change source information; recovery checks availability rather than semantic fidelity. The downstream evaluation uses recovered text while retaining the original construction status. The registered budget band tests the text, not the instruction. For a source with N context tokens, a requested ratio r sets the integer target T = ⌊rN ⌋. Let R be the token count of the decoded relational text under the tokenizer used for that request. The registered acceptance band is 0.98T ≤ R ≤ T. (1) The upper bound enforces the ceiling; the lower bound tests whether generation fills the requested allowance. These conditions distinguish controlling the amount of delivered text from merely staying below a maximum. The count excludes envelope syntax, question text, and reader instructions. After an initial attempt, the wrapper permits one revision using token-count feedback. Only 33 of 9,600 budgeted outputs satisfy this band. Figure 1 shows where the parsed lengths fall across the requested grid. Both anchors produce median lengths above their requests throughout that grid. The response to a budget instruction thus preserves a substantial gap between the intended allowance and the delivered context. The plotted lengths use the Qwen tokenizer and the last parsed attempt. Bars describe variation across outputs, rather than uncertainty about a mean. Unparsed outputs remain in the closure denominator even though they supply no length to the plot. Likewise, recovered bodies can contribute lengths while their raw envelopes remain schema violations. Delivered tokens determine the reading comparison. The realized ratio associates the returned text with its measured context cost under each consumer’s tokenizer. A plot against requested ratios instead describes how the procedure responds to different instructions. The next section uses measured ratios to place the matrix’s answer scores alongside delivered context length. 4 R EADING ACCURACY AT DELIVERED LENGTHS The natural-length matrix resolves encoder and reader choices. Table 1 reports answer scores for each encoder–consumer pair, with full prose as the common reference. 3
Text of page 4
Realized token ratio GPT-5.6 Sol Sonnet 5 1.1 1.1 0.8 0.8 0.5 0.5 0.2 0.2 0.25 0.4 0.55 0.7 0.25 Requested token ratio Parsed n: 1184, 1198, 1200, 1200 identity 0.4 0.55 0.7 Requested token ratio Parsed n: 1200, 1198, 1200, 1198 Last parsed outputs; bars show 10th–90th percentiles. Figure 1: Do requested budgets predict realized lengths? Among parsed outputs, medians and 10th–90th percentiles exceed the identity line on the Qwen tokenizer; 33 of all 9,600 cells close within 98–100% of target. Table 1: Descriptive natural-length modified token-overlap F1 by encoder and consumer, with full prose as the reference. Entries pool examples; construction failures receive zero F1. Encoder Llama F1 Mistral F1 Qwen F1 GPT-5.6-sol Llama 4 Maverick Magistral Small Ministral 8B Sonnet 5 0.3954 0.3501 0.4024 0.4009 0.3796 0.3909 0.3741 0.3754 0.3781 0.3834 0.5378 0.4252 0.5220 0.4897 0.4824 Full prose 0.4066 0.3679 0.5621 The direction relative to full prose differs across consumers. Every relational-text entry in the Mistral column lies above its full-prose reference; the entries in the Llama and Qwen columns lie below theirs. The encoder ordering also changes with the consumer. These observations make the pair, rather than either model in isolation, the unit of the descriptive comparison. The dataset-balanced paired contrast is −0.0251, with a 95% interval from −0.0356 to −0.0146. This summary first compares methods on the same questions, then equally weights datasets, encoders, and consumers. The table instead pools examples within each pair, preserving the composition of the fixed sample. Appendix A gives the paired calculation and resampling procedure. A separate consumer-input measure gives a pooled TE-to-prose token ratio of 0.8875. It includes the question and reader template and equally weights datasets, encoders, and consumers. The contextonly realized ratio isolates the rewritten part of the input; the complete-prompt ratio describes its contribution to the full reader input. Construction outcomes are part of the measured procedure. Sonnet’s natural-length condition lacks usable text on 87 of 1,200 rows, and one GPT budgeted condition fails on 16 of 1,200 rows. 4
Text of page 5
Both fall below the registered construction threshold. The analysis retains these failures as empty answers and restricts the comparisons to descriptive reporting. Realized lengths locate the available accuracy–cost observations. Figure 2 places answer F1 against measured context length for the two budgeted anchors, with a separate panel for each consumer. Each point equally weights the dataset means, unlike the example pooling in Table 1. Lines follow the sequence of budget requests and show how the delivered inputs change along that sequence. Successful constructions contribute their measured context-token ratios. Failures have no realized readable input, so the registered placement uses the requested position and F1 zero. The resulting aggregates include both outcomes under this explicit placement rule. Modified F1 Llama 3.1 Qwen 3.5 0.6 0.6 0.5 0.5 0.5 0.4 0.4 0.4 0.3 0.3 0.3 Mistral 7B 0.6 0.7 1.1 0.3 0.3 0.7 1.1 0.3 0.7 1.1 Mean token ratio (measured; requested on failure) GPT-5.6 Sol Sonnet 5 Figure 2: Dataset-balanced modified token-overlap F1 across budget requests. Horizontal means use measured context-token ratios for successful constructions and requested ratios for failures; failures also receive zero F1. Lines connect the recorded points for each encoder. Comparator curves are unavailable. An equal-cost comparison would require comparator outcomes at corresponding delivered lengths. Comparator purpose determines the inference. Full prose answers how a constructed input compares with the original supplied context. The registered direct-prose rewrite instead controls for encoding the same source under a length constraint, making it the required counterpart for the planned form comparison. Its development requirement failed; the pruning baseline had no matrix consumer pass; a second baseline had unresolved usage rights. The planned targeted-prompt calibration was also absent. Consequently, every registered confirmatory comparison is unavailable. Appendix D links each intended comparison to its missing prerequisite. 5 F ROM TEACHER TARGETS TO AN OPEN COMPILER The compiler study asks whether a trained open model can supply the relational representation through an explicit plan. Construction proceeds from source units to propositions with source references, then from that plan to readable TE. These stages provide distinct observations: the plan can be checked mechanically, its serialized text can be counted, and the consumer can be evaluated on that text. Learning from grouped source units. The open Qwen3-Next-80B teacher produces plans for individual source units, which are assembled into groups before a separate call supplies dependency links. The training set contains 27,284 accepted group examples. A Qwen3-8B student learns 5
Text of page 6
question-agnostic and question-conditioned plan construction; the acceptance test uses the questionagnostic output (TE-Cache). The compiler supplies relations and source references, and a fixed serializer turns them into readable Telegraph English. The plan represents propositions and their grounding separately. A span table names intervals in the source units; proposition rows carry predicates, arguments, qualifiers, and references to those spans. Dependency rows link propositions through supported relations such as coreference, temporal precedence, or qualification. The prompt asks for preservation of negation, modality, attribution, and scope because these features can change the content a reader receives even when the main entities remain. It also keeps answer labels and downstream scores outside the plan-construction input. Accepted groups are filtered by source and sequence limits. The training construction caps the source at 7,500 characters and 20 units, and excludes sequences beyond the training limit. Evaluation sources extend to 24,000 characters, making longer contexts a test beyond that construction envelope. The full-parameter student is trained on the accepted groups with both plan modes; Appendix F records target assembly, the earlier target diagnosis, training, and filtering rules. Training runs for two epochs with an 8,192-token maximum sequence length. Only output tokens receive labels, and question-agnostic and question-conditioned modes have equal target-token weight. The two modes share one fitted model; the acceptance comparison uses the question-agnostic compilation from that model. A plan allowance and a text budget control different stages. The proposition allowance grows with source-unit count and character length under a fixed rule shared by training and evaluation. That rule also determines the visible-token allowance for the emitted JSON plan. It gives longer and more divided sources room to express more propositions, up to the registered cap. The consumer’s requested compression ratio enters afterward, when the completed plan is rendered as TE. For K source units and L source characters, the implemented rule is M = clamp(max(K + 2, round(L/350)), 4, 32), m = min(K, M ), T plan = 150M + 40K. The prompt uses m, M for its proposition allowance and T plan for its target-token allowance. This rule depends on source size, whereas the later reader-text budget depends on the requested ratio and consumer tokenizer. The serializer treats the requested token budget as a ceiling. The registered rule selects the longest fitting rendering among fixed-content punctuation variants and records failure when none fits. It neither deletes propositions nor pads a short representation to meet the target. This construction can satisfy a ceiling while leaving substantial budget unused. For a fixed plan, increasing the ceiling expands the eligible renderings without rerunning content selection. The resulting text is therefore determined jointly by what the compiler put in the plan and which registered rendering fits the consumer tokenizer. Writing V(p) for the fixed-content renderings of plan p and τ c for the consumer’s token count expresses the registered selection rule: z ∗ ∈ arg max τ c (z). z∈V(p): τ c (z)≤B The serializer fails if no rendering fits the ceiling B. Increasing B changes this feasible set, while the propositions in p remain fixed. Thus a sweep over reader budgets measures how the existing plan is delivered; it does not ask the compiler for progressively more source facts. The comparison uses a budget-targeted LongLLMLingua baseline, with realized-budget equivalence unverified in the final record. Plan validity is measurable before answer accuracy. Table 2 reports valid compilations across the evaluation sources. Each question is compiled in both training modes, so the denominator counts plans rather than distinct questions. The runtime check requires a JSON root object with the expected schema, source hash, and mode, nonempty proposition and span fields, and a proposition count within the cap. These checks establish basic construction integrity; complete span-boundary and semantic validation are separate requirements. The in-domain sources yield mostly valid plans, while the longer QASPER inputs frequently exhaust the output ceiling. After compilation, the serializer can still reject a malformed plan row or a plan 6
Text of page 7
Table 2: Valid compiler plans by dataset, pooling question-agnostic and question-conditioned modes. The out-of-envelope QASPER inputs have substantially fewer valid plans. Dataset Valid plans Total plans 599 593 1,157 488 600 600 1,200 1,000 2WikiMultiHopQA HotpotQA MuSiQue QASPER whose text does not fit the requested ceiling. Answer evaluation measures the text that reaches the consumer after these stages, retaining construction failures on questions with answers as zero F1. The compiler fails the measured acceptance criteria. Figure 3 shows the compiler and pruning baseline across the requested budget grid. The width-normalized area is 0.0658 for the compiler and 0.3936 for LongLLMLingua on the in-domain datasets. The difference is −0.328, with a 95% interval from −0.349 to −0.307. None of the three consumer families meets the required positive-contrast criterion. This measured contrast fails the registered acceptance rule. Modified F1 From source to answers 0.5 Source units 0.4 Open compiler plan 0.3 Source-sized plan allowance 0.2 Fixed-content serializer 0.1 Requested text ceiling Longest fitting variant Reject if none fits 0 0.25 0.4 0.55 0.7 Requested token ratio Delivered TE Pooled across in-domain datasets and consumer families Consumer answers Open compiler (TE-Cache) LongLLMLingua (budget-targeted) Figure 3: Modified token-overlap F1 across requested budgets for the question-agnostic compiler and the budget-targeted pruning baseline. The schematic locates the plan allowance and consumer-text ceiling in the compression procedure. The requests do not establish equal realized token counts. Appendix G reports the complete acceptance decision and unmeasured components. The curve evaluates responses to the registered budget requests. At each request, its score averages the available in-domain datasets and consumer families; the width-normalized area summarizes those points under the frozen scoring rule. For requested ratios r 1 < · · · < r 4 and those mean scores f (r j ), 7
Text of page 8
the compiler area is A = 3 X 1 f (r j ) + f (r j+1 ) (r j+1 − r j ). r 4 − r 1 j=1 2 Its horizontal intervals come from the requested grid, whereas the matrix’s raw area uses measuredratio positions with the stated failure convention. The baseline rises across the displayed requests, while the compiler’s scores remain low. Plan construction, the fixed-content serializer, and the reader all lie on the path measured by this contrast. Two components remain unmeasured. The terminal answer file has no answers for the 500 QASPER questions, a pre-existing data defect that makes the out-of-distribution accuracy component unscorable. These cases are excluded and counted separately from construction failures on questions with answers, which receive F1 zero. Citation-reach judging was prepared but never executed. The acceptance summary in Appendix G distinguishes these missing evaluations from the measured failures. The cost component first requires a requested-budget point where compiler accuracy matches the comparator. No point meets that prerequisite, so the study has no eligible accuracy-matched point for its reuse-cost comparison. The retained observations connect mostly valid in-domain plans to poor answer scores after serialization and reading. Locating the lost evidence within that path requires the generated plan, delivered text and measured length, and paired reader outcome. 6 R ELATED WORK What reaches the reader. Readable compression changes the text supplied to a language model. Selective Context removes lexical units using self-information, while LLMLingua allocates budgets and prunes tokens iteratively (Li et al., 2023; Jiang et al., 2023a). LongLLMLingua incorporates question relevance and information positioning; LLMLingua-2 learns token selection from distilled data (Jiang et al., 2024; Pan et al., 2024). RECOMP trains extractive and abstractive compressors for downstream task performance (Xu et al., 2024). Gist tokens, AutoCompressors, and the In-context Autoencoder compress context into learned activations or memory slots (Mu et al., 2023; Chevalier et al., 2023; Ge et al., 2024). Token selection can directly change how much source material is retained, whereas rewriting generates a new string whose length depends on its wording. These different construction operations motivate observing each compressor’s output before treating a requested rate as a shared condition. Our setting uses inspectable text between a distinct encoder and consumer, making the delivered string and its token count observable parts of the interface. Relational form and its construction. Telegraph English rewrites prose as compact relational statements; subsequent work compares this form with token-matched controls and post-truncated prose summaries (Arbuzov et al., 2026; Bei et al., 2026). StructGPT supplies linearized evidence from structured sources, while PICARD enforces output syntax during decoding (Jiang et al., 2023b; Scholak et al., 2021). These approaches connect a representational choice to a procedure that produces usable inputs. The present evaluation follows that connection through schema compliance, mechanically recovered text, measured length, and consumer answers. Its matrix tests several encoder–consumer pairings under fixed construction rules. Learning a compressor from targets. Sequence-level distillation trains a student on teacher outputs, and rationale distillation adds intermediate explanations (Kim & Rush, 2016; Hsieh et al., 2023). RECOMP likewise trains a smaller compressor from generated summaries selected using downstream signals (Xu et al., 2024). The targets determine which output structures the student is taught to produce. Our compiler study examines source-unit grouping and output allowances in those targets, then evaluates the learned plan after fixed serialization. Plan validity and downstream answer quality provide separate measurements of this construction-and-reading procedure. Which comparison a score answers. Performance-versus-budget reporting exposes resource dependence (Dodge et al., 2019); long-context studies distinguish available information from its effective use (Liu et al., 2024). LLMLingua discusses generated-length control and reports actual token use alongside accuracy (Jiang et al., 2023a). Pre-registration fixes intended comparisons before 8
Text of page 9
outcomes are observed (van Miltenburg et al., 2021). Building on these practices, we distinguish requested budgets, realized inputs, and the evidence needed for a representation comparison. Full prose anchors the original context; an encoder-produced prose rewrite supplies a different control. Their separate roles determine what the available answer scores can establish. 7 D ISCUSSION The matrix demonstrates that a common requested ratio does not place encoders at a common delivered length. The compiler shows the complementary limit: enforcing a text ceiling is compatible with poor answer quality. Measure the input the consumer receives. A budget instruction specifies a desired output; the delivered representation determines the reader’s task. Relational text can be recovered from a malformed wrapper while exceeding its requested length. A serializer can respect a ceiling while leaving much of the allowance unused. These cases explain why schema compliance, text availability, and token use need separate measurements. Retaining the actual string and its consumer-specific token count makes answer quality interpretable at the input that produced it. The tokenizer and inclusion of prompt overhead also belong with the reported ratio. Choose the control for the scientific question. Full prose asks how a constructed representation performs relative to the original supplied context. A prose rewrite from the same encoder at a comparable delivered length addresses the additional question of representational form. The matrix provides the former comparison descriptively; its registered controlled comparison remains unavailable. The compiler supplies a measured failure of its acceptance rule over requested budgets. That outcome evaluates the trained compiler, serializer, and consumers as a procedure responding to budget requests. Comparison at equal realized lengths requires the corresponding measured inputs from both arms. Connect construction quality to answer quality. At evaluation, a valid plan establishes the implemented syntactic and reference checks; consumer answers measure the usefulness of its serialized text. Keeping construction failures in the scored denominator evaluates the complete procedure, including its ability to deliver an input. Questions without answer keys instead define an unavailable evaluation component. Across the fixed English question-answering data and tested model versions, the studies make the path from compression request to answer score explicit. Their practical consequence is a comparison record that retains the request, returned text, measured length, construction outcome, and paired alternative. That record makes delivered evidence, rather than the requested allowance alone, the basis of the comparison. 9
Text of page 10
R EPRODUCIBILITY STATEMENT The supplement contains the available numerical exports, protocols, recovered analysis sources, figure scripts, and an inventory of missing replay inputs. It supports rebuilding the reported displays; complete model-level replay additionally requires the unavailable raw examples, predictions, and trained weights. Timestamps are withheld from the anonymous version for restoration at cameraready; the artifact chain records their order. E THICS STATEMENT The study evaluates context compression on existing question-answering datasets. Its principal resource consideration is the compute used for teacher generation, compiler training, and consumer inference. Cost reporting covers the recorded component measurements. AI U SE S TATEMENT Generative models also produced teacher targets for compiler distillation, as described in Section 5. In this work, generative AI tools were used to polish the manuscript prose and assist with preparing explanatory illustrations. Generative AI tools were not used to design the model. All AI-assisted text and visual materials were reviewed and edited by the authors. The authors take responsibility for the final content of this paper, including all text, claims, and artifacts produced with AI assistance. R EFERENCES Mikhail L. Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, and Alexey A. Shvets. Telegraph English: Semantic prompt compression via structured symbolic rewriting, 2026. URL https: //arxiv.org/abs/2605.04426v1. Version 1. Sisong Bei, Mikhail L. Arbuzov, Ziwei Dong, Dmitri Kalaev, and Alexey Shvets. Context compression is not one thing: Readable symbolic re-expression vs. coherent summary at matched budget, 2026. URL https://arxiv.org/abs/2606.14875v1. Version 1. Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. Adapting language models to compress contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 3829–3846. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.232. URL https://aclanthology.org/2023.emnl p-main.232/. Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4599–4610. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.naacl-main.365. URL https://aclanthology.org/2021.naac l-main.365/. Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. Show your work: Improved reporting of experimental results. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pp. 2185–2194. Association for Computational Linguistics, 2019. doi: 10.18653/v1/D19-1224. URL https://aclanthology.org/D19-1224/. Tao Ge, Jing Hu, Lei Wang, Xun Wang, Si-Qing Chen, and Furu Wei. In-context autoencoder for context compression in a large language model. In International Conference on Learning Representations, 2024. URL https://proceedings.iclr.cc/paper_files/pape r/2024/hash/0b276510ec2d3f6613a8b60c41ff0438-Abstract-Conference. html. Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. Constructing a multihop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th 10
Text of page 11
International Conference on Computational Linguistics, pp. 6609–6625. International Committee on Computational Linguistics, 2020. doi: 10.18653/v1/2020.coling-main.580. URL https: //aclanthology.org/2020.coling-main.580/. Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 8003–8017. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.findings-acl.507. URL https://aclanthology.org/2023. findings-acl.507/. Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 13358–13376. Association for Computational Linguistics, 2023a. doi: 10.18653/v1/2023.emnlp-main.825. URL https://aclanthology.org/2023.emnlp-main.825/. Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1658–1677. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.acl-long.91. URL https://aclanthology.org/2024.ac l-long.91/. Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Xin Zhao, and Ji-Rong Wen. StructGPT: A general framework for large language model to reason over structured data. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 9237–9251. Association for Computational Linguistics, 2023b. doi: 10.18653/v1/2023.emnlp-main.574. URL https://aclanthology.org/2023.emnlp-main.574/. Yoon Kim and Alexander M. Rush. Sequence-level knowledge distillation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1317–1327. Association for Computational Linguistics, 2016. doi: 10.18653/v1/D16-1139. URL https: //aclanthology.org/D16-1139/. Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. Compressing context to enhance inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 6342–6353. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.391. URL https://aclanthology.org/2023.em nlp-main.391/. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2024. doi: 10.1162/tacl_a_00638. URL https://aclanthology.org/2024.tacl-1.9/. Jesse Mu, Xiang Lisa Li, and Noah Goodman. Learning to compress prompts with gist tokens. In Advances in Neural Information Processing Systems, volume 36, 2023. doi: 10.52202/075280-0848. URL https://proceedings.neurips.cc/paper_files/paper/2023/hash/3 d77c6dcc7f143aa2154e7f4d5e22d68-Abstract.html. Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. LLMLingua- 2: Data distillation for efficient and faithful task-agnostic prompt compression. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 963–981. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.findings-acl.57. URL https: //aclanthology.org/2024.findings-acl.57/. Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 9895–9901. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.emnlp-main.779. URL https://aclanthology.org/2021.emnlp-main.779/. 11
Text of page 12
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. MuSiQue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics, 10:539–554, 2022. doi: 10.1162/tacl_a_00475. URL https: //aclanthology.org/2022.tacl-1.31/. Emiel van Miltenburg, Chris van der Lee, and Emiel Krahmer. Preregistering NLP research. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 613–623. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.naacl-main.51. URL https: //aclanthology.org/2021.naacl-main.51/. Fangyuan Xu, Weijia Shi, and Eunsol Choi. RECOMP: Improving retrieval-augmented LMs with compression and selective augmentation. In International Conference on Learning Representations, 2024. URL https://proceedings.iclr.cc/paper_files/paper/2024/file/ bda88ed2892f5e61c9a9bf215c566913-Paper-Conference.pdf. Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2369–2380, 2018. doi: 10.18653/v1/D18-1259. URL https://aclantholo gy.org/D18-1259/. 12
Text of page 13
A
S TATISTICAL PROCEDURE AND SCORE DEFINITIONS
A.1
F ROZEN ANSWER SCORE
The recorded scorer normalizes prediction and answer strings by lowercasing, removing articles,
removing non-word punctuation, and collapsing whitespace. Let P and G be the resulting token
lists, and let u = | set(P ) ∩ set(G)|. Precision is u/|P | and recall is u/|G|; their harmonic mean is
the recorded answer F1. The maximum score over the supplied answer aliases is used. An empty
normalized prediction receives zero against a nonempty normalized gold answer; the primitive
returns one when both normalized lists are empty. The compiler wrapper separately forces empty or
non-complete predictions to zero when gold answers exist.
This implementation uses distinct overlapping tokens in the numerator and full list lengths in the
denominators. It therefore differs from ordinary multiset token F1 when a word repeats. The same
frozen function is used by both studies. The reported scores preserve that implementation; no
benchmark rescoring was performed for this manuscript.
A.2
M ATRIX STATISTICS
The report sample consists of 600 MuSiQue, 300 2WikiMultiHopQA, and 300 HotpotQA questions.
Table 1 averages all 1,200 rows separately for each encoder–consumer pair. The shared full-prose
observation is evaluated once per question and consumer and joined by identifier when forming
paired contrasts. This join preserves the question as the resampling unit.
For each dataset, encoder, and consumer, the descriptive contrast averages the paired TE-minus-full-prose score differences. The pooled contrast equally averages those 45 cell means. The frozen
bootstrap draws 10,000 multinomial question-count vectors within each dataset using NumPy PCG64
and seed zero. Each draw applies the same counts to all paired methods, encoders, and consumers.
Intervals use linearly interpolated 2.5th and 97.5th percentiles. The descriptive interval is unadjusted
and carries no confirmatory test interpretation.
The pre-registered confirmatory procedure included family-level testing and Benjamini–Hochberg
correction. Every registered test is unavailable because a required condition is absent or ineligible.
The export assigns these tests a bookkeeping value of p = 1.0; this value is a status convention rather
than an observed test result. Appendix D reports the statuses directly.
A.3
T OKEN ACCOUNTING AND MATRIX FRONTIERS
For question i and consumer c, let N ic be the token count of its full-prose context. The requested
budget is T icr = ⌊rN ic ⌋, with r ∈ {0.25, 0.40, 0.55, 0.70}. The generation wrapper accepts counts
in [0.98T icr , T icr ] and permits one count-feedback revision. Figure 1 uses saved medians and
percentiles of the last parsed attempts under the Qwen tokenizer. Unparsed outputs are absent from
those quantiles but remain in the closure denominator.
The matrix frontier counts the decoded representation string with the named consumer tokenizer and
no added special tokens. JSON syntax, question text, and reader instructions are excluded. Successful
constructions contribute their measured representation-to-prose ratio; failures with no realized count
contribute their requested position and F1 zero. Each Figure 2 point equally averages three dataset
means from the supplied pooled display. Lines connect observed aggregates; no intermediate budget
outcomes were generated.
The registered matrix area is the raw trapezoidal integral over mean token-ratio positions at the
compressed grid points, using measured lengths for successes and requested positions for failures.
It is computed per dataset, anchor, and consumer before the registered aggregation. The full-prose
reference is excluded from this integral. Comparator areas are unavailable, so this paper displays the
available curves without a comparator-ranking claim.
The separate consumer-input ratio includes the complete rendered prompt, including question and
reader template. Its supplied pooled value, 0.8875, averages the 45 dataset–encoder–consumer
means. The exported per-cell dictionary drops the dataset key and retains only MuSiQue entries
after overwriting. Accordingly, the manuscript uses the code-verified pooled scalar and omits that
13
Text of page 14
misleading per-cell table. The original export and implementation are preserved in the supplement
with this defect documented.
A.4
C OMPILER STATISTICS
The compiler scorer shares the answer-F1 and paired-bootstrap primitives above. At each requested
ratio, it equally averages F1 over datasets with available answers and over consumers. If f (r) denotes
this mean, its area is
X f (r j ) + f (r j+1 )
1
A =
(r j+1 − r j ).
r max − r min j
2
Here the abscissae are the requested ratios, not achieved token counts. This width-normalized score
differs from the matrix’s raw token-ratio integral.
The contrast is the question-agnostic compiler area minus the available LongLLMLingua area. The
same within-dataset question resample applies across arms, consumers, and budgets. Only records
marked complete supply predictions; other record states receive zero when gold answers exist.
QASPER records have empty answer lists and are excluded and counted separately. Their absence
leaves the out-of-domain acceptance clause unscorable.
B
M ODELS , PROMPTS , AND CONSTRUCTION
B.1
M ATRIX IMPLEMENTATION
The primary TE prompt requests one fact per line, flat readable relations, factual preservation,
unambiguous abbreviations, and a fixed JSON envelope. It is question-agnostic, task-informed, and
tokenizer-neutral. Its output schema is {"curated_te": "...", "issues": []}. The
token count covers the decoded relational text, excluding envelope keys and escapes. Mechanical
recovery extracts text from recognized surface forms without using answer outcomes. The primary
analysis retains recovery version 2; a later recovery revision is outside the primary result.
Direct budget-conditioned prose was intended to use the same encoder and source, targeting the TE
output’s consumer-token budget. It allowed count-only revisions without truncation or answer-based
selection. Its development failure prevents this paired representation comparison. LongLLMLingua
was intended to receive source and question; BRIEF-Pro required development-only calibration of its
native sentence-count control. The required comparison outputs were unavailable for the matrix.
The five encoder families are fixed model points across four provider lineages, with Magistral and
Ministral sharing a provider. The secondary Ministral size panel expands the schema census but not
the primary family average. Consumer weights and tokenizers are identified by immutable revisions
in the supplement. The frozen manifest records the shared prose and matrix readers’ matching code
and prompt hashes.
The recovered reader implementation uses each model’s chat template, greedy generation, and a
maximum of 48 new tokens. It disables optional thinking, checks for an open reasoning block, and
records over-window inputs instead of truncating them. The exact external reader prompt remains
hash-identified but absent from the supplied files. Recovered worker versions are preserved as archival
implementation evidence rather than asserted to be identical to every earlier executed reader binary.
B.2
S OURCE SAMPLING
The fixed report sample uses deterministic hash ordering within the registered hop strata. MuSiQue
contains 311 two-hop, 189 three-hop, and 100 four-hop questions. 2WikiMultiHopQA contains 100
questions in each of those strata; HotpotQA contains 300 two-hop questions. Development examples
are separate and serve engineering and calibration decisions. The compiler evaluation adds 500
QASPER contexts to these 1,200 in-domain examples. Raw selected contexts and answer files are
not contained in the supplied archive.
14
Text of page 15
C The broader generation census separates correctly formed outputs from text recovered mechanically after a schema violation. Table 3 copies the generated census counts. The correctly formed column requires a nonempty relational-text field; Sonnet also has eight well-formed but empty outputs. Recovery availability is a text-construction measure, not a validation of semantic fidelity. S CHEMA COMPLIANCE AND RECOVERED TEXT Table 3: Generation census across 1,200 inputs per encoder. Correctly formed nonempty outputs and mechanically recovered nonempty text measure different construction properties. Encoder Correct form Usable text 1,200 0 2 0 0 0 202 1,200 1,200 1,200 1,200 1,182 1,200 1,113 GPT-5.6-sol Llama 4 Maverick Magistral Small Ministral 14B Ministral 3B Ministral 8B Sonnet 5 The matrix nevertheless contains all 50,400 expected reading records across its 42 condition cells. Record presence and usable construction are separate checks. Sonnet’s natural condition has construction completeness 0.9275, with a dataset minimum of 0.8983. The GPT budgeted condition identified in the main text has completeness 0.9867. The primary analysis retains failed constructions as empty answers. D The registered matrix analysis contains eleven unavailable comparisons. Table 4 groups them by the missing prerequisite while retaining the registered identifiers. Every listed comparison has status unavailable; no entry represents an executed null test. Table 4: Registered matrix comparisons and their missing prerequisites. All listed tests are unavailable. U NAVAILABLE MATRIX COMPARISONS Registered comparison Missing prerequisite H1; H2a; H2b; H4a Direct-prose construction failed development; dependent form comparisons unavailable. Usage rights unresolved. No matrix consumer pass. Targeted-prompt calibration not generated. Direct-prose comparator unavailable. BRIEF-Pro comparator unavailable. Pruning comparator unavailable. H3 vs. BRIEF-Pro H3 vs. LongLLMLingua H4b; H4c Frontier vs. direct prose Frontier vs. BRIEF-Pro Frontier vs. LongLLMLingua H1 concerns the pooled representation contrast; H2a and H2b concern the specified encoder conditions and their contrast. H4a measures form, H4b prompt adaptation, and H4c the consumer interaction in adaptation. E A VAILABLE FRONTIER POINTS Table 5 reproduces the supplied dataset-balanced means used in Figure 2. The underlying per-dataset table and source summaries accompany the figure script. Failure placement and pooling follow Appendix A. 15
Text of page 16
Table 5: Dataset-balanced frontier points by encoder, consumer, and requested ratio. The token-ratio column uses measured lengths for successes and requested positions for construction failures; failures receive zero modified token-overlap F1. Encoder Consumer GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol GPT-5.6-sol Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Sonnet 5 Llama Llama Llama Llama Mistral Mistral Mistral Mistral Qwen Qwen Qwen Qwen Llama Llama Llama Llama Mistral Mistral Mistral Mistral Qwen Qwen Qwen Qwen Requested Token ratio F1 0.25 0.4 0.55 0.7 0.25 0.4 0.55 0.7 0.25 0.4 0.55 0.7 0.25 0.4 0.55 0.7 0.25 0.4 0.55 0.7 0.25 0.4 0.55 0.7 0.3338 0.6060 0.7562 0.8452 0.3566 0.6462 0.8031 0.8952 0.3496 0.6339 0.7880 0.8805 0.6512 0.8677 0.9894 1.0671 0.6708 0.8849 1.0038 1.0792 0.6574 0.8670 0.9838 1.0585 0.4153 0.4326 0.4535 0.4598 0.3760 0.4272 0.4427 0.4416 0.4266 0.5217 0.5284 0.5507 0.4558 0.4761 0.4769 0.4842 0.4484 0.4558 0.4590 0.4517 0.5403 0.5518 0.5674 0.5597 16
Text of page 17
F C OMPILER TARGETS , TRAINING , AND VALIDATION F.1 F IRST - COMPILER DIAGNOSIS The primary diagnosis artifact measures 29,634 teacher targets, of which 14,634 are question-agnostic and 15,000 are question-conditioned. Every target has one source unit; 84.32% reach the sevenproposition cap. The Spearman correlation between proposition count and source length is 0.2988. Only 827 targets carry a dependency, and 73.26% place all cited spans in the first half of the source. The median maximum cited-span endpoint, divided by source length, is 0.432. These quantities describe training targets. The second compiler changes target grouping, prompt and schema, assembly, and the training length distribution together. These joint changes leave the contribution of each target property unresolved. F.2 S ECOND - COMPILER CONSTRUCTION The open teacher is Qwen3-Next-80B and the student is Qwen3-8B. The accepted training set contains 27,284 groups, comprising 15,000 question-agnostic and 12,284 question-conditioned examples. Per-unit teacher outputs are validated, assembled with a one-to-one unit mapping, linked through a separate dependency call, and validated again at group level. Dependency absence is recorded rather than rejected. Assembly relabels span and proposition identifiers while preserving sourceunit mappings and remapping within-unit links. The separate dependency call receives proposition identifiers, predicates, argument surfaces, and information about units from the same document; it can add links across unit boundaries. The student receives the grouped source and produces its corresponding plan in one compilation. Let K be the source-unit count and L its character count. The output allowance is set by M = clamp(max(K + 2, round(L/350)), 4, 32), m = min(K, M ), T = 150M + 40K. Here M and m specify the proposition allowance and T the target-token allowance used by the prompt construction. The same rule is applied to evaluation inputs, and the supplement includes the exact implementation. Training construction enforces a source ceiling of 7,500 characters and at most 20 units. Evaluation sources extend to 24,000 characters. The registered group-drop allowance is at most 25%; sequenceover-limit examples are excluded and logged. The student uses full-parameter bfloat16 training with full-shard FSDP, activation checkpointing, and no packing. Maximum sequence length is 8,192; AdamW uses learning rate 2 × 10 −5 , weight decay 0.1, and gradient clipping at 1.0. Training lasts two epochs with cosine scheduling and 3% warmup. Only output tokens receive labels, and the two modes have equal target-token weight. The effective batch allowance is 131,072 tokens across eight GPUs, with micro-batch one and accumulation two. The configuration records separate data-order, initialization, and dropout seeds; these are components of one fitted model, not independent training replicates. The archived recipe notes approximately 2.5× more target tokens per row than for the first compiler at the same learning rate. F.3 W HAT A VALID COMPILATION CHECKS The recovered worker checks a JSON root object, source hash, mode, nonempty propositions, proposition cap, and a nonempty span table. Its valid-plan count does not check every span boundary or entailment relation. The counts in Table 2 measure this implemented validation over both compilation modes. The 2Wiki summary is recovered verbatim from an embedded audit input; the other dataset summaries are standalone files. MuSiQue has 41 JSON-invalid plans and two source-hash failures. QASPER has 496 JSON-invalid plans at the output ceiling, eight excessive-proposition plans, and eight prompt-over-context failures. These leave 488 valid plans among its 1,000 mode-specific attempts. The archived HotpotQA coverage field is 1.0002, which exceeds a bounded fraction; it is not used as a coverage result here. QASPER’s separate coverage field is conditional on valid plans and is not a whole-sample success rate. 17
Text of page 18
G The pre-registered acceptance rule requires all components to pass. Table 6 states their final dispositions under the terminal scorer. The measured failures determine the overall failed decision despite the unavailable components. Table 6: Compiler acceptance components, final status, and the supporting observation. Incomplete and unavailable components remain distinct from measured failures. C OMPILER ACCEPTANCE AND RECORD COMPLETENESS Component Status Observation Pooled area Failed Paired interval Families and transfer Failed Incomplete Citation reach Open compiler Unavailable Failed Matched-accuracy cost Failed Compiler area is below the budget-targeted comparator. The required positive lower endpoint is absent. Zero of three families pass; QASPER answers are missing. Judge tasks were prepared but never submitted. The compiler does not pass the pooled-area and interval requirements. No requested-budget point meets the prerequisite accuracy comparison. The family component requires a positive lower endpoint of the 95% interval for at least two of three consumers and for the out-of-domain contrast. The missing QASPER answers prevent evaluation of its latter clause. The cost component tests measured reuse costs only at accuracy-matched points; its failure identifies the missing accuracy prerequisite rather than a measured cost disadvantage. The final consumer ledger contains 66,300 records: 57,910 complete, 8,375 construction failures, 14 refusals, and one over-context record. These counts span full prose, both compiler modes, the pruning baseline, and all datasets. The 19,500 QASPER records are excluded from F1 because their answer lists are empty. All other failed record states remain in the paired scored set with zero F1. The historical protocol requested a comparator-crash sensitivity analysis, but that output is absent from the supplied final artifacts. The terminal paired score retains comparator failures under its zero-score rule. Available component timing and cost records are provided without an amortized cost-saving conclusion. H P ROTOCOL ORDER AND ARTIFACT AVAILABILITY The supplied protocols distinguish decisions made before answer outcomes from the terminal analysis. Matrix order. First, the roster, sampling, neutral prompt, comparisons, and analysis contract were fixed. Engineering checks then determined comparator eligibility and mechanical text recovery. After generation, completeness was checked while answer outcomes remained sealed. A recorded decision authorized descriptive reading with construction failures retained. The terminal analysis then reported the descriptive outcomes and unavailable confirmatory tests. Compiler order. The protocol first specified its open-compiler construction and acceptance conditions. An open-teacher amendment fixed the training-source boundary. A serializer amendment changed exact equality to a token ceiling after a check on training targets, before method-arm evaluation outputs existed. The second training recipe and grouped-target construction were fixed before that training run. Compilation, serialization, and consumer records preceded the terminal acceptance score. Rebuildable artifacts. The reviewer supplement contains final numerical summaries, figure input extracts and scripts, protocol copies, receipt-pinned analysis sources, and recovered compiler and reader workers. It includes an extraction manifest for source files embedded in audit requests. 18
Text of page 19
Declared release transforms remove calendar timestamps and infrastructure identifiers from reviewer copies while preserving scientific values and code behavior. Model revisions and scientific configuration identifiers remain available where needed to identify an implementation. Missing replay inputs. Raw question contexts, answer files, individual predictions, full teacher targets, and fitted compiler weights are absent from the supplied handoff. Some exact prompt files and imported rendering or serialization modules are also missing. Consequently, the supplement rebuilds figures and permits inspection of the saved statistics and available implementations; it cannot rerun the complete experiments from this archive alone. 19