What Survives Learned Symbolic Compression?
Back to the paper page. ICLR 2027 submission, August 2026.
All 22 pages are shown below.
Text of page 1
W HAT S URVIVES L EARNED S YMBOLIC C OMPRESSION ? Anonymous authors Paper under double-blind review A BSTRACT Lossy text compression for language-model pipelines is judged by reconstruction distance or by downstream task accuracy, and neither says whether the compressed representation still states what the source stated. We measure agreement with a constructed reference for grammar-constrained symbolic codes. The reference represents source facts as equations, and a fixed decoder recovers the equations a code encodes. An exact rule system compares their consequences independently of the code’s well-formedness checker. On controlled arithmetic micro-worlds, we evaluate learned translators from 70M to 2.8B parameters against structured controls and learned baselines at matched byte and token budgets. Passing the checker is not preserving the source: a self-consistent mistranslation passes it, and so do sampled learned-translation errors. Across byte budgets, the best learned trace translator trails the strongest compact structured control, with area under the closure-agreement curve of 0.67 against 0.86. Increasing translator size does not close this aggregate gap; the largest Pythia translator matches the best smaller translator’s saved byte-axis agreement scores. Source fidelity relative to the constructed reference, internal validity, and what a consumer model recovers are three different quantities and should be measured separately. 1 I NTRODUCTION A compressed representation becomes a substitute for its source. Its usefulness depends on what survives that substitution, including facts that a downstream task may never ask about. Prompt pruning and learned summaries reduce the cost of long contexts, but successful downstream answers test only the information those answers require (Jiang et al., 2023; Pan et al., 2024; Xu et al., 2024). We study a complementary question: how much of a specified set of source facts and consequences survives in the representation itself? Symbolic codes make this question concrete because their statements can be decoded and checked. Yet a checker can accept a code that describes the wrong world. Suppose the source states x = 2, y = x + 3, and z = 2y. A code that changes the offset to y = x + 4 can consistently compute and check z = 12. Its internal checks succeed even though the source entails z = 10. Distinguishing these outcomes requires a comparison with the source, independent of the code’s own checks. We make that comparison exact on controlled affine arithmetic. A reference projection records source facts as canonical rational equations, and a fixed decoder maps the code into the same equation language. Exact rules form finite sets of explicit equations, entailed values, and pairwise differences. Their Jaccard overlap is closure agreement; a trace accepted independently by the code’s checker is linter-valid. A separate consumer task measures whether another model recovers the named values from the representation. The construction connects source-grounded compression with closure-based distortion for deductive sources (Trukhina & Vashkelis, 2026a; Xu, 2026c;a). These instruments expose different properties of learned compression. Sampled learned outputs contain valid codes with source errors. Across byte budgets, the strongest compact structured control retains more measured content than the best trace translator; larger translators leave this aggregate gap open in the evaluated training setup. Token accounting changes the leading structured control. Consumer recovery then tests the representations with a model that must extract their values, under fixed response conditions. Reviewers: please read the Reviewer Guidelines (iclr.cc/Conferences/2027/ReviewerGuidelines) and the AI Policy for Reviewers (iclr.cc/Conferences/2027/AIPolicyForReviewers). If you used AI to expand, edit, or polish your review, please provide the input text to the LLM. Better still, consider skipping the LLM and submitting your original text: we, and the authors, are much more interested in your unedited thoughts than in what an LLM has to say. AI-assisted or not, you are putting your name and reputation behind your review: LLM-generated falsehoods, hallucinations or misrepresentations are subject to disciplinary action, which may include desk-rejecting all papers you have authored. 1
Text of page 2
2 The measurement compares two descriptions of the same world: the source-side reference and the facts extracted from a compressed code. Figure 1 follows this comparison through the running example. The checker has a separate role: it tests the code’s grammar and supported internal constraints using only the code. M EASURING WHAT SURVIVES THE SUBSTITUTION From text to comparable facts. The reference projection is determined before any compressed code exists. It contains the explicit equations associated with the rendered source, including deliberately redundant statements. Canonicalization combines coefficients, reduces rational values, and normalizes equation order and sign. Thus z = 2y and 2y − z = 0 identify the same explicit equation. The encoder receives the rendered text; the construction record supplies the reference used for scoring. Each supported code format has a fixed decoder into this equation language. For trace code, the decoder reads supported GIVEN bindings and affine EQ equalities. Free-text reasoning and other trace fields remain outside the extracted fact set, while unsupported material remains visible in parsing diagnostics. Canonical JSON, compact structured codes, and template prose have their own fixed parsers. This shared output language permits comparisons across formats without requiring a common surface syntax. Constructed source x is 2; y is x plus 3; z is 2 times y. Reference projection Self-consistent mistranslation x=2 y=x+3 z=2*y GIVEN: x := 2 GOAL: z EQ[e1]: y = x + 4 EQ[e2]: z = 2*y CHECK[c1]: arith: e2.value == 12 ANS: 12 Apply fixed consequence rules Code-only checker: accepts Decode; apply the same rules Reference closure Decoded-code closure 2*y-z=0; x-y=-3 x-z=-8; x=2 y-z=-5; y=5; z=10 2*y-z=0; x-y=-4 x-z=-10; x=2 y-z=-6; y=6; z=12 Source agreement: 2/12 = 0.167 Exact-code fixture: checker accepts; closure agreement 1.000 Figure 1: Can valid code change its source? This constructed fixture changes x + 3 to x + 4 and passes its code-only checker, but closure agreement falls from the exact-code fixture’s 1.000 to 0.167. A finite consequence signature. Write C(S) for the set containing explicit canonical equations, uniquely entailed variable values, and uniquely entailed pairwise differences from facts S. Exact rational row reduction tests these consequences. For source facts R and decoded code facts Z, agreement is |C(R) ∩ C(Z)| J(R, Z) = . |C(R) ∪ C(Z)| An empty decoded set or an inconsistent side receives zero agreement. The variable domain includes variables on either side, making both missing and added facts visible. Retaining explicit equations also means that equivalent affine bases can receive different scores when their explicit members 2
Text of page 3
fall outside the value and difference families. This finite signature defines the preservation target throughout the paper. Using a set makes repeated copies of the same consequence count once. Using the union in the denominator also distinguishes agreement from recall alone: extra decoded facts can reduce the score even when the reference facts remain recoverable. Within this signature, each distinct equation, value, or difference contributes equally. The resulting score measures overlap in the chosen consequences, rather than the number of source sentences copied or the importance of a fact to a particular downstream question. What the example measures. The reference in Figure 1 has seven distinct consequences under this rule. They include the values x = 2, y = 5, and z = 10, their pairwise differences, and the relation 2y − z = 0. An exact code preserves all seven and scores one. The consistent mistranslation preserves x = 2 and 2y − z = 0, but changes the other values and differences. Two shared consequences among twelve in the union give agreement 1/6. Checking its answer against its own equations accepts the changed world; comparison with the reference identifies the change. The same example distinguishes useful omission from lost information. One supplied fixture adds the explicit source statement z = 10. Omitting that statement from the code preserves agreement one because x = 2, y = x + 3, and z = 2y still entail it. In another fixture, omitting z = 2y removes the connection to z and leaves only three of the seven reference consequences, giving agreement 3/7. These are constructed fixtures that explain the instrument; learned-output evidence follows in Section 4. The comparison also accounts for additions. A further supplied fixture keeps every reference equation and adds w = 9, a value absent from the source. The original seven consequences remain, but the decoded signature grows to eleven, giving agreement 7/11. The added value and its differences with the existing quantities enter the denominator because the variable domain includes both descriptions. Recovering every source fact therefore suffices for full recall but not full agreement. Content under an output allowance. We evaluate nominal fractions 0.30, 0.45, 0.60, and 0.80 of each world’s nonredundant source rendering. Absolute allowances have floors of 128 bytes or 48 reference tokens. Using the same nonredundant source to set allowances prevents added redundant text from granting extra output space. Byte and token conditions generate separate outputs under their respective allowances. Raw output accounting includes whitespace, fences, explanations, and malformed material. An over-budget code stays in the evaluation denominator and receives zero matched-budget agreement. For mean agreement F̄ j at ordered allowances r j , closure-AUC is AUC = 3 X F̄ j + F̄ j+1 1 (r j+1 − r j ) . 0.80 − 0.30 j=1 2 Every budget point contributes to this normalized area; the curve retains the operating points that the aggregate combines. Appendix A specifies world-level aggregation and resampling. Bytes and reference tokens are separate cost units, so we retain both axes when comparing formats. The redundancy analysis makes a further controlled comparison: it adds already entailed values and differences while preserving the reference consequence set. Code sizes are compared only for pairs meeting the registered agreement and checker requirements on both renderings. 3 E XPERIMENTAL DESIGN AND THE ROLES OF THE CONTROLS The sources describe controlled micro-worlds whose named quantities are linked by exact affine relations: sums, constant multiples, and offsets. An anchored, connected construction determines rational values, expressed through varied sentences and statement order. All renderings of a world stay together in one split. Training uses 24,000 worlds; held-out reporting uses 200 worlds with 800 renderings, and consumer evaluation uses a fixed subset of 60 worlds with 240 renderings. Separate development and selection worlds support implementation choices and checkpoint selection. Appendix A.7 specifies the splits, adapters, selection rule, and available provenance. The construction supplies both training targets and evaluation references. Every fidelity score compares the decoded equations with that constructed reference. Exact affine equalities provide a 3
Text of page 4
controlled setting for tracing information through compression. The reference language excludes negation, modality, quantifiers, time, and uncertainty. Learned encoders with fixed decoders. Trace translators use low-rank adapters on Pythia checkpoints from 70M to 2.8B parameters, with three seeds labelled A, B, and C (Biderman et al., 2023). Additional Qwen translators provide a second-family comparison (Yang et al., 2025). The completed grid contains 39 training cells across translators and learned controls. Adapter rank 16 and learning rate 0.0002 were chosen on development data. All cells use an effective batch size of 128 and at most three epochs; checkpoint selection maximizes closure-AUC on the separate selection worlds, choosing the earliest checkpoint in a tie. Thus the size comparison evaluates the codes produced by a common training and selection procedure. The structured controls—canonical JSON, CCL-Core, CCL-Min, and fixed-template prose—use trained Pythia 1.4B encoders and deterministic decoders. CCL-Core and CCL-Min adapt source-grounded formats to this domain (Trukhina & Vashkelis, 2026a). These comparisons evaluate complete encoding systems, including learning errors and representation cost. The structured controls vary the target format; the scale ladder keeps the trace format fixed while varying the translator. The release supports aggregate comparisons but omits raw learned outputs, weights, and the imported construction and decoding implementations, including the adapted CCL serialization (Appendix A.7). Construction and task controls. The irredundant-core oracle reads the construction directly and serializes its correct core facts. It shows how a complete reference code performs under the same output allowances. LLMLingua-2 pruning and SemanticZip-style learned or strong encoders supply task-oriented comparisons (Pan et al., 2024; Trukhina & Vashkelis, 2026b). These arithmetictask adaptations receive cost and consumer scores; their formats have no supplied deterministic equation decoder for primary closure scoring. Appendix H records the separate unconstrained-prose construction attempt. Recovering values from representations. Consumers receive the representation and canonical entity identifiers, then return a value map of integers or reduced fractions. The consumers are Qwen3-Next-80B and GPT-OSS-20B, with a registered 256-token response cap (Qwen Team, 2025). A secondary GPT-OSS condition keeps its prompt and changes decoding settings. Each output scores the fraction of requested values recovered exactly, with parsing failures counted as incorrect. The reported score averages these fractions. Source agreement, validity, and recovery thus use separate instruments, with the saved statistical procedures specified in Appendix A. 4 P ASSING THE CHECKER LEAVES SOURCE ERRORS UNDETECTED Learned translators also produce internally valid codes that change their source. Some sampled codes satisfy both their checker and their output allowance while losing source agreement. Table 1 summarizes the saved diagnostic samples. For each translator, the analysis selects 40 failing output records by a deterministic ordering of item hashes. Failure here means imperfect budgeted agreement, which includes both source disagreement and exceeding the allowance. The table therefore retains the over-budget counts alongside checker acceptance. Table 1: Checker outcomes among 40 sampled failing output records per translator. Failures include source disagreement and budget violations; the two columns can overlap. Translator Linter-valid Over budget Pythia 410M Pythia 1B Pythia 1.4B Pythia 2.8B 8 11 12 14 7 0 5 0 Qwen 0.6B Qwen 1.7B 16 15 0 1 4
Text of page 5
The samples contain 8–16 linter-valid outputs per model. The Pythia 1B, Pythia 2.8B, and Qwen 0.6B samples contain no over-budget records; their accepted failures therefore have imperfect source agreement. These rows establish the source-error case directly. For rows with budget violations, the marginal counts leave their overlap with checker acceptance unresolved. The diagnostic samples identify failure modes rather than estimate their prevalence across all outputs. Their selection conditions on failure, and an item can contribute output records under different evaluation conditions. Appendix F gives the sampling rule and complete taxonomy. We next apply that source-conditioned comparison to every output, including failures, to measure retention across budgets. 5 A COMPACT STRUCTURED CONTROL RETAINS MORE CONTENT PER BYTE The strongest trace translator has lower byte-axis closure-AUC than the strongest structured control: 0.672 against 0.860 for CCL-Min. Figure 2 shows how the aggregate difference arises across the available budgets. The comparison asks how much measured content each complete encoding system delivers within a shared allowance. It includes the translator’s choice of facts, the cost of expressing them, and whether the resulting code fits. Mean closure agreement CCL-Min CCL-Core Structured prose 1.0 0.5 0.0 0.30 0.45 0.60 0.80 0.30 Canonical JSON 0.45 0.60 Trace, 1B 0.80 0.30 0.45 0.60 0.80 Trace, 160M 1.0 0.5 0.0 0.30 0.45 0.60 0.80 0.30 0.45 0.60 0.80 0.30 0.45 0.60 0.80 Nominal byte budget / nonredundant source size Figure 2: How much closure agreement fits each byte budget? Shared axes show each system’s four saved measurements, including JSON’s initial decline and the low 160M trace scores. Where the byte advantage appears. The displayed 1B trace system retains less measured content than CCL-Min under the tighter allowances and approaches it at the largest allowance. The displayed trace series is the best aggregate performer; individual budget points can favor other translators. The saved trace scores rise from 0.281 and 0.514 at the two tighter byte allowances to 0.789 and 0.995 at the larger ones. CCL-Min gives 0.459, 0.811, 0.994, and 0.999 at the same operating points. At allowance 0.60, CCL-Min already approaches complete agreement while the trace system retains less of the measured content. 5
Text of page 6
The comparison also distinguishes structured formats from one another. Canonical JSON reaches byte-axis closure-AUC 0.624, below the best trace result. CCL-Min’s aggregate advantage therefore concerns the strongest compact format, rather than every structured representation. The saved JSON curve also declines between its first two allowances before rising. Each operating point uses a separate generated output, rather than extending the same code from the preceding point. Accordingly, the curves report observed retention under each allowance; they need not rise monotonically. Changing the unit changes the comparison. Token accounting changes the ordering among structured controls. Canonical JSON leads that axis at 0.514, compared with 0.426 for CCL-Min; trace translators reach at most 0.369. Table 2 places the structured controls’ two aggregates together. Table 2: The leading structured control changes with the cost unit. Each column aggregates outputs generated for that budget axis; it is not a recount of one shared set of codes. Structured control Byte closure-AUC Token closure-AUC Canonical JSON CCL-Core CCL-Min 0.624 0.776 0.860 0.514 0.471 0.426 Appendix B reports both axes for all systems. A byte-efficient representation need not lead under the reference tokenizer: the budget unit changes both the generated outputs and the observed ranking. Correct facts still have a representation cost. The irredundant-core oracle separates exact source access from compact representation. Its byte-axis agreement is zero at 0.30 and perfect from 0.60, when its complete code fits. This code reads the construction record directly and retains its irredundant equations. Even with direct access to correct facts, their chosen serialization must fit the allowance. These budget comparisons establish an aggregate advantage for the compact structured control. The next question is whether increasing the trace translator’s capacity closes that advantage under the same measurement. 6 L ARGER TRANSLATORS REACH A MEASURED PLATEAU Increasing translator size improves the weakest trace systems but does not close the aggregate gap to CCL-Min. Figure 3 places each saved Pythia seed and the Qwen summaries against the structured reference. The comparison keeps the trace representation and evaluation target fixed while varying the translator within the model-size ladder. The 70M and 160M systems remain near zero, while the 410M translator reaches 0.654. The 1B and 2.8B models both reach 0.672, and the intervening 1.4B model reaches 0.668. The Qwen summaries are 0.586 and 0.587. The larger Pythia systems share a plateau in byte-axis closure agreement. Moving beyond the weakest models brings a large change in measured retention; subsequent capacity increases yield a much narrower range of byte-axis scores. The recorded byte-axis curves for 1B and 2.8B coincide, while token-axis and validity measurements differ. The plateau is therefore a result about saved agreement scores under this training setup. A plateau and an equivalence decision ask different questions. The scale decision combines an AUC equivalence margin of 0.02 with a bound on the difference in checker pass rates. No smaller Pythia model meets that combined criterion. The 1B and 1.4B comparisons satisfy the AUC interval condition, but their checker-rate differences exceed the joint rule’s allowance. For 410M, the AUC interval extends past the equivalence boundary. Appendix D reports both quantities. 6
Text of page 7
Byte closure-AUC All translators Larger Pythia: detail CCL-Min: 0.860 0.8 0.67 0.6 0.4 0.65 0.2 0.0 0.63 70M 160M 410M 1B 2.8B 410M 1B 1.4B 2.8B Translator parameters (log scale in each panel) Pythia mean Qwen3 mean Seeds A / B / C: -4 / 0 / +4 pt horizontal offsets Figure 3: Does scale close the measured byte-axis gap? Scores stay below CCL-Min; the larger-Pythia detail resolves seed variation, with circles showing seeds A/B/C from left to right. Table 3: Why similar byte-axis AUC does not satisfy the joint scale rule. Differences are from Pythia 2.8B. AUC intervals must lie inside (−0.02, 0.02); checker pass-rate differences must have magnitude at most 0.02. Pythia model AUC difference: 95% interval Checker-rate difference [−0.021, −0.015] same AUC [−0.005, −0.002] −0.084 −0.049 −0.042 410M 1B 1.4B Removing redundant wording is a separate property. The redundancy test reports a zero mean log-size change in each evaluated translator-seed comparison. This aggregate stability is consistent with the specified training target, a canonical representation shared by redundant renderings. The test compares each rendering at its own first qualifying budget cell, where its source agreement and checker result meet the registered requirements. For nonredundant and redundant sources x 0 , x 3 with qualifying codes z 0 ∗ , z 3 ∗ , it measures ∆ log B = log B(z 3 ∗ ) , B(z 0 ∗ ) ∆ log r = log B(z 3 ∗ )/B(x 3 ) . B(z 0 ∗ )/B(x 0 ) The first quantity tracks absolute code size; the second tracks size relative to the expanded source. Evaluated comparisons contain 3–27 qualifying world pairs. Systems with too few qualifying pairs have undefined tests, as retained in Appendix C. The result characterizes stability against redundant source wording among qualifying pairs. It can coexist with the remaining gap in agreement across the full budget range, where other outputs lose content or fail to fit. The budget and scale results measure content through a fixed decoder. The final comparison replaces that decoder-based question with a practical one: can another model use the representation to recover the named values? 7
Text of page 8
7 Qwen3-Next-80B recovers values more accurately from canonical JSON than from CCL-Min in the consumer experiment. Recovery spans both budget axes on a report subset; the headline closure-AUC uses the byte axis on the full report set. Table 4 compares exact value recovery under the supplied consumer protocol. For each output, the score is the fraction of requested values returned correctly; the table averages these fractions. C ONSUMER RECOVERY UNDER FIXED RESPONSE CONDITIONS Table 4: Exact value recovery by Qwen3-Next-80B. Source and oracle inputs provide reference points; structured and task-oriented encodings show recovery from compressed representations. Input Exact recovery Source text Irredundant-core oracle 0.869 0.907 Canonical JSON Trace, 1B translator CCL-Min 0.703 0.586 0.471 SemanticZip-style adapter Strong-encoder SemanticZip-style LLMLingua-2 0.463 0.787 0.169 Canonical JSON and the displayed 1B trace code exceed CCL-Min’s recovery point estimate for this consumer. This is a comparison of recovery under the consumer protocol; relating its ordering to closure agreement would require both measurements on the same codes and worlds. The consumer’s recovery point estimate is higher for the oracle’s explicit equations than for the rendered source. The source score is thus a reference-input measurement, not an upper bound on recovery. The strong-encoder SemanticZip-style encoding provides the highest displayed recovery among those task-oriented encodings. Primary GPT-OSS-20B recovery rounds to zero for all but the strong-encoder input, with generation frequently reaching the 256-token cap. The secondary condition keeps the prompt and token cap fixed and changes the reasoning-effort setting. It recovers some values across systems, showing that the consumer measurement also depends on decoding conditions. Appendix E gives every condition and its recorded stop reasons. These condition-specific point estimates describe what each consumer recovers; Appendix A documents the separate pooled comparison procedure. 8 R ELATED WORK Compression objectives. Prompt compression reduces the material presented to a language model. Selective Context and the LLMLingua family remove tokens, with objectives ranging from information selection to question relevance and task-agnostic retention (Li et al., 2023; Jiang et al., 2023; 2024; Pan et al., 2024). RECOMP learns summaries of retrieved passages; gist tokens and AutoCompressors learn continuous representations for subsequent model use (Xu et al., 2024; Mu et al., 2023; Chevalier et al., 2023). Telegraph English uses compact symbolic fact lines, and related experiments compare re-expression with summaries and token-matched controls (Arbuzov et al., 2026; Bei et al., 2026). These approaches motivate measuring both representation cost and consumer behavior. Our fixed decoder supplies an additional comparison against specified source consequences. Preservation beyond task success. Context Codec evaluates typed facts and their source support, while SemanticZip distinguishes protected from lossy content and measures recovery by another model (Trukhina & Vashkelis, 2026a;b). The controls here adapt these ideas to affine sources. Closure-based rate–distortion theory supplies the precedent for comparing consequence sets, including deductive sources and reversible logging (Xu, 2026c;a;b). We instantiate that pattern with a finite equation signature and learned text-to-code translators. Semantic communication also studies meaning 8
Text of page 9
preservation under communication constraints, using learned encoders and decoders (Xie et al., 2021). Here exact decoding makes the scored fact sets inspectable. Faithfulness and executable meaning. Summarization studies identify unsupported generated content; FactCC predicts source-conditioned consistency, and FActScore evaluates support at the level of atomic facts (Maynez et al., 2020; Kryściński et al., 2020; Min et al., 2023). The constructed arithmetic reference permits an exact comparison of both missing and added consequences. Semantic parsing links language to executable representations, while constrained decoding enforces structural requirements (Liang et al., 2013; Scholak et al., 2021). Our separate checker and reference comparison distinguish these structural properties from preservation of the source. Learned representations and scale. CommNet, DIAL, referential games, and GLC study messages through the behavior they enable, including connections to interpretable symbols (Sukhbaatar et al., 2016; Foerster et al., 2016; Lazaridou et al., 2017; Du et al., 2026). Consumer recovery retains that behavioral question alongside source agreement. Language-model scaling studies and the Pythia family motivate systematic size comparisons (Kaplan et al., 2020; Hoffmann et al., 2022; Biderman et al., 2023). Our size ladder evaluates the resulting codes under fixed preservation rules and output allowances, connecting translator capacity with the content retained by a complete encoding system. 9 D ISCUSSION The learned-output errors and budget curves identify two requirements for a useful substitute: preserving the specified source consequences and expressing them within the output allowance. Choose what must survive. The reference projection turns preservation into an explicit set comparison. The fixtures distinguish dropping a redundant statement from losing a relation, and a self-consistent mistranslation exposes the information missing from a code-only check. Learned outputs exhibit the same separation between internal acceptance and source agreement. For a pipeline that replaces its source with code, the preservation target therefore belongs in the evaluation design alongside the code grammar. Choose the cost that matters. The byte and token results lead to different choices among structured controls. CCL-Min leads aggregate agreement per byte, while JSON leads under token accounting. These are properties of complete encoding systems: a trained translator must recover the source facts and express them within the allowance. The oracle makes the latter requirement visible even with direct access to correct facts. An encoding comparison becomes useful when its budget matches the resource the intended pipeline must conserve. Evaluate the substitute where it will be used. Translator size, checker acceptance, and consumer recovery provide complementary evidence about a code. The Pythia byte-axis plateau coexists with differences in checker pass rates, so replacing a translator on the basis of one score can change another property. The consumer experiment adds the model that must interpret the representation and the decoding procedure that produces its answer. Its descriptive ordering supplies a separate view of the compressed inputs, with the evaluation population and response conditions stated alongside the scores. The exact arithmetic domain makes this separation inspectable. Extending it to open text requires a source representation and consequence family appropriate to that domain; the finite signature already makes those choices part of the present score. The source-to-code comparison identifies changed content; the consumer test measures what a particular reader can recover from the resulting representation. 9
Text of page 10
R EPRODUCIBILITY S TATEMENT The materials include measurement contracts, score tables, display scripts and data, configuration and checkpoint hashes, and the analysis implementation. Seeds are labelled A, B, and C in the paper; their integer values remain in the supplied configuration. Timestamps are withheld from the anonymous version and restored at camera-ready; the artifact chain fixes the order. Appendix A.7 lists the additional artifacts required for complete replay. E THICS S TATEMENT The study uses constructed arithmetic sources and a restricted exact-value consumer task. Its central safety implication is that a code’s internal acceptance should remain distinct from evidence of faithful translation. AI U SE S TATEMENT In this work, generative AI tools were used to polish the manuscript prose and assist with preparing explanatory illustrations. Generative AI tools were not used to design the model. All AI-assisted text and visual materials were reviewed and edited by the authors. The authors take responsibility for the final content of this paper, including all text, claims, and artifacts produced with AI assistance. R EFERENCES Mikhail L. Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, and Alexey A. Shvets. Telegraph English: Semantic prompt compression via structured symbolic rewriting, 2026. URL https: //arxiv.org/abs/2605.04426v1. Version 1. Sisong Bei, Mikhail L. Arbuzov, Ziwei Dong, Dmitri Kalaev, and Alexey Shvets. Context compression is not one thing: Readable symbolic re-expression vs. coherent summary at matched budget, 2026. URL https://arxiv.org/abs/2606.14875v1. Version 1. Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. Pythia: A suite for analyzing large language models across training and scaling. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 2397–2430. PMLR, 2023. URL https://proceedings.mlr.press/v202/biderman23a.html. Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. Adapting language models to compress contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 3829–3846. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.232. URL https://aclanthology.org/2023. emnlp-main.232/. Wei Du, Benyu Wu, Yuqing Sun, Wei Guo, Yuntao Du, Zhongmin Yan, Guoxian Yu, and Lizhen Cui. Learning efficient and interpretable multi-agent communication. In International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id= a3CUE06G5Y. B. Efron. Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7(1):1–26, 1979. doi: 10.1214/aos/1176344552. URL https://www.jstor.org/stable/2958830. Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, and Shimon Whiteson. Learning to communicate with deep multi-agent reinforcement learning. In Advances in Neural Information Processing Systems, volume 29, 2016. URL https://proceedings.neurips.cc/paper_files/ paper/2016/file/c7635bfd99248a2cdef8249ef7bfbef4-Paper.pdf. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom 10
Text of page 11
Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. Training compute-optimal large language models, 2022. URL https://arxiv.org/abs/ 2203.15556. Sture Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2):65–70, 1979. URL https://www.jstor.org/stable/4615733. Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 13358–13376. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.825. URL https://aclanthology.org/2023.emnlp-main.825/. Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1658–1677. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.acl-long.91. URL https://aclanthology.org/2024. acl-long.91/. Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models, 2020. URL https://arxiv.org/abs/2001.08361. Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9332–9346. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.emnlp-main.750. URL https://aclanthology.org/2020.emnlp-main.750/. Daniël Lakens. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science, 8(4):355–362, 2017. doi: 10.1177/1948550617697177. URL https://journals.sagepub.com/doi/10.1177/1948550617697177. Angeliki Lazaridou, Alexander Peysakhovich, and Marco Baroni. Multi-agent cooperation and the emergence of (natural) language. In International Conference on Learning Representations, 2017. URL https://arxiv.org/abs/1612.07182v2. Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. Compressing context to enhance inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 6342–6353. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.391. URL https://aclanthology.org/2023. emnlp-main.391/. Percy Liang, Michael I. Jordan, and Dan Klein. Learning dependency-based compositional semantics. Computational Linguistics, 39(2):389–446, 2013. doi: 10.1162/COLI_a_00127. URL https: //aclanthology.org/J13-2005/. Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 1906–1919. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.acl-main.173. URL https://aclanthology.org/2020.acl-main. 173/. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 12076–12100. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-main.741. URL https://aclanthology. org/2023.emnlp-main.741/. 11
Text of page 12
Jesse Mu, Xiang Lisa Li, and Noah Goodman. Learning to compress prompts with gist tokens. In Advances in Neural Information Processing Systems, volume 36, 2023. doi: 10.52202/075280-0848. URL https://proceedings.neurips.cc/paper_files/ paper/2023/hash/3d77c6dcc7f143aa2154e7f4d5e22d68-Abstract.html. Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven, and David T. Mellor. The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11):2600–2606, 2018. doi: 10.1073/pnas.1708274114. URL https://www.pnas.org/doi/10.1073/pnas. 1708274114. Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 963–981. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.findings-acl.57. URL https://aclanthology. org/2024.findings-acl.57/. Qwen Team. Qwen3-Next-80B-A3B-Instruct. Model card, 2025. URL https://huggingface. co/Qwen/Qwen3-Next-80B-A3B-Instruct. Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 9895–9901. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.emnlp-main.779. URL https://aclanthology.org/2021.emnlp-main.779/. Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus. Learning multiagent communication with backpropagation. In Advances in Neural Information Processing Systems, volume 29, 2016. URL https://papers.nips.cc/paper_files/paper/2016/hash/ 55b1927fdafef39c48e5b73b5d61ea60-Abstract.html. Natalia Trukhina and Vadim Vashkelis. Compress the context, keep the commitments: A formal framework for verifiable LLM context compression, 2026a. URL https://arxiv.org/abs/ 2605.17304v1. Version 1. Natalia Trukhina and Vadim Vashkelis. SemanticZip: A pilot framework for lossy text compression with LLMs as semantic decompressors, 2026b. URL https://arxiv.org/abs/2605. 24541v1. Version 1. Huiqiang Xie, Zhijin Qin, Geoffrey Ye Li, and Biing-Hwang Juang. Deep learning enabled semantic communication systems. IEEE Transactions on Signal Processing, 2021. doi: 10.1109/TSP.2021. 3071210. URL https://arxiv.org/abs/2006.10685v3. Fangyuan Xu, Weijia Shi, and Eunsol Choi. RECOMP: Improving retrieval-augmented LMs with compression and selective augmentation. In International Conference on Learning Representations, 2024. URL https://proceedings.iclr.cc/paper_files/paper/2024/file/ bda88ed2892f5e61c9a9bf215c566913-Paper-Conference.pdf. Jianfeng Xu. Rate-distortion theory for deductive sources under closure fidelity, 2026a. URL https://arxiv.org/abs/2604.15698v4. Version 4. Jianfeng Xu. Closure-preserving rate-distortion for reversible logging, 2026b. URL https: //arxiv.org/abs/2606.16592v2. Version 2. Jianfeng Xu. Semantic rate-distortion theory: Deductive compression and closure fidelity, 2026c. URL https://arxiv.org/abs/2604.11204v1. Version 1. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang 12
Text of page 13
Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388. 13
Text of page 14
A
S TATISTICAL PROCEDURE AND IMPLEMENTATION
A.1
R EFERENCE FACTS AND FINITE CLOSURE
Canonical equations combine repeated variables, sort identifiers, clear denominators, divide by the
positive greatest common divisor, and orient the first nonzero coefficient positively. Exact rational
row reduction tests entailment over the union of source and decoded variables.
For a consistent equation set S, the scored set contains its explicit canonical equations, all uniquely
entailed entity values, and all uniquely entailed pairwise differences. Explicit affine equations remain
members even when they fall outside the value and difference templates. Consequently, logically
equivalent affine bases can have different scored sets.
Redundant source additions are restricted to entailed values and pairwise differences, keeping the
scored reference set fixed across the four source renderings. An inconsistent side, or an empty
decoded set, receives zero agreement. Unsupported text remains visible in parsing diagnostics.
A.2
R ATES AND AGGREGATION
Let B count UTF-8 bytes and T count reference tokens. Raw outputs retain whitespace, fences,
explanations, and malformed material. The reference tokenizer is Qwen/Qwen3.8-27B, revision
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, with special tokens disabled. For the
rendering x 0 without added redundancy, the absolute allowances are
b B (r) = max{128, ⌊rB(x 0 )⌋},
r ∈ {0.30, 0.45, 0.60, 0.80}.
b T (r) = max{48, ⌊rT (x 0 )⌋},
(1)
(2)
Realized rates use the corresponding source rendering in the denominator. An over-budget output
receives zero matched-budget agreement while its raw output and diagnostics remain recorded. Each
budget cell contributes an item-level agreement score.
For mean scores F̄ j at ordered budgets r j , normalized closure-AUC is
AUC =
3
X
F̄ j + F̄ j+1
1
(r j+1 − r j )
.
0.80 − 0.30 j=1
2
(3)
The analysis requires every budget point. Scale summaries retain separate seed results and their
equal-weight mean.
A.3
W ORLD RESAMPLING
The resampling unit is the latent world. All associated renderings, budgets, systems, and seeds move
together. The percentile interval uses the ordered replicate positions at 0.025 and 0.975, with linear
interpolation between neighboring positions (Efron, 1979).
Scale contrasts and the designated Pythia 1.4B redundancy analysis use 10,000 resamples. The allmodel redundancy summaries and baseline contrasts use 1,000 resamples. The numerical bootstrap
seed is retained in the configuration.
Scale and redundancy computations retain the multiplicity of sampled worlds. Baseline contrasts first
collapse records by item identifier, so repeated selections of a world contribute once per replicate.
These baseline intervals therefore use the implemented deduplicated-world procedure.
A.4
R EDUNDANCY COMPARISON
The first passing budget requires closure precision and recall of at least 0.95 and a full-linter pass. A
rendering without a passing cell receives ordinal code 5. Size and rate comparisons use only worlds
14
Text of page 15
passing in both the R0 and R3 renderings. Each rendering uses its own first passing cell: B(z 3 ∗ ) , B(z 0 ∗ ) B(z 3 ∗ )/B(x 3 ) ∆ log r = log , B(z 0 ∗ )/B(x 0 ) ∆ pass = Pr(pass at 0.80 | R3) − Pr(pass at 0.80 | R0). ∆ log B = log (4) (5) (6) The rule requires the upper interval bound for ∆ log B to be at most log(1.05) and the upper bound for ∆ log r to be negative. It also requires the lower bound for ∆ pass to be at least −0.02. Maximumbudget success is evaluated specifically at 0.80, independently of any earlier success. A.5 S CALE COMPARISON The reference is Pythia 2.8B. A smaller model qualifies only when its paired AUC interval lies strictly inside (−0.02, 0.02) and its full-linter pass-rate difference has magnitude at most 0.02. An adjacent larger-minus-smaller contrast with upper interval bound below −0.02 prevents a scale threshold. An equivalence margin expresses a chosen tolerance for differences (Lakens, 2017). The Qwen models remain separate diagnostic points. The reported scale analysis consists of seed means, paired reference contrasts, and adjacent contrasts. The proposed mixed-model slope, secondary token-scale contrasts, and redundancy sensitivity regression have no corresponding outputs in the analysis record. A.6 B ASELINE AND CONSUMER COMPARISONS The comparison adapter is Pythia 1.4B, as specified before evaluation. Pythia 1B supplies the separate best-byte-AUC descriptive comparison. Each baseline contrast averages paired item differences after averaging repeated records for the same item. The implemented task contrast first averages every stored consumer score per output, including the secondary decode control. The consumer tables retain the three conditions separately. Thus the pooled contrast and each named consumer’s recovery score are different summaries. The implementation assigns each baseline an indicator of 1.0 if any paired interval includes zero, and 0.01 otherwise. Holm adjustment operates across baselines within each metric family, pooling the byte and token cells for that indicator (Holm, 1979). These inputs are interval-derived indicators, rather than calibrated hypothesis-test p-values. The recorded dominance flag requires nonnegative lower bounds in every cell, positive lower bounds in at least two cells, and rejection after the indicator adjustment. The archived analysis retains these flags; they are not used as hypothesis-test evidence in this paper. The original protocol instead specified separate axes and consumers, with adjustment across operating points. A.7 D ATA AND TRAINING SETTINGS Training uses 24,000 worlds; development and selection each use 80 worlds and 320 renderings. The held-out report set contains 200 worlds and 800 renderings. The consumer subset contains the 60 smallest-hash report worlds and all 240 renderings. Every rendering of a world remains in its assigned split. LoRA rank 16 and learning rate 0.0002 were selected on development data before checkpoint selection. Training uses alpha twice the rank, dropout 0.05, bfloat16, effective batch size 128, and at most three epochs. AdamW settings are (β 1 , β 2 ) = (0.9, 0.95), epsilon 10 −8 , and weight decay 0.1. The schedule is cosine, with warmup ratio 0.03 and gradient clipping 1.0. Attention and MLP projections receive adapters. Checkpoint selection maximizes selection closure- AUC, choosing the earliest checkpoint within 10 −6 ties. Training seeds are A, B, and C; the release configuration maps these labels to integers. 15
Text of page 16
The checkpoint record includes matched Pythia 1.4B adapters for JSON, CCL-Core, CCL-Min, structured prose, and SemanticZip-style packets. The structured formats have deterministic decoders. The aggregate tables identify the output format of each control. Each translator produces one greedy completion with temperature zero, repetition penalty one, and EOS enabled. The generation allowance adds 32 tokens for byte-controlled cells and 8 for token-controlled cells. All returned tokens remain in rate accounting. Available artifacts comprise aggregate score tables, the analysis driver, and configuration and checkpoint hashes. Raw scorecards, learned output texts, fitted weights, and the imported estimator, decoder, target-builder, and construction implementations are absent from this record. The exact adapted CCL serialization is therefore unavailable for independent inspection. Aggregate results can be inspected, while complete row-to-checkpoint attribution and raw-output equality remain unverified. B R ATE – DISTORTION RESULTS The complete scored-format results retain both budget axes. The source, pruning, and packet controls receive consumer scores because they lack deterministic equation decoders. Table 5: Mean matched-budget closure agreement and normalized AUC on the bytes axis. System 0.30 0.45 0.60 0.80 AUC TE, 1.4B TE, 160M TE, 1B TE, 2.8B TE, 410M TE, 70M TE, Qwen3 0.6B TE, Qwen3 1.7B Canonical JSON CCL-Core CCL-Min Oracle core Structured prose 0.276 0.008 0.281 0.281 0.276 0.005 0.233 0.233 0.277 0.337 0.459 0.000 0.324 0.509 0.008 0.514 0.514 0.510 0.005 0.417 0.417 0.236 0.648 0.811 0.725 0.613 0.788 0.008 0.789 0.789 0.771 0.005 0.648 0.648 0.890 0.948 0.994 1.000 0.905 0.992 0.008 0.995 0.995 0.949 0.005 0.999 1.000 1.000 0.999 0.999 1.000 1.000 0.668 0.008 0.672 0.672 0.654 0.005 0.586 0.587 0.624 0.776 0.860 0.768 0.749 Table 6: Mean matched-budget closure agreement and normalized AUC on the tokens axis. System 0.30 0.45 0.60 0.80 AUC TE, 1.4B TE, 160M TE, 1B TE, 2.8B TE, 410M TE, 70M TE, Qwen3 0.6B TE, Qwen3 1.7B Canonical JSON CCL-Core 0.068 0.001 0.090 0.090 0.071 0.002 0.088 0.090 0.150 0.132 0.146 0.003 0.207 0.210 0.142 0.003 0.201 0.202 0.325 0.290 0.269 0.005 0.383 0.393 0.270 0.004 0.386 0.384 0.562 0.524 0.539 0.007 0.775 0.774 0.566 0.005 0.772 0.769 0.988 0.902 0.256 0.004 0.365 0.369 0.261 0.004 0.363 0.362 0.514 0.471 16
Text of page 17
Continued from previous page C CCL-Min Oracle core Structured prose 0.113 0.230 0.403 0.994 0.426 0.000 0.000 0.005 1.000 0.202 0.143 0.279 0.453 0.781 0.420 0.60 0.80 AUC R EDUNDANCY RESULTS ∆ log B Model Seed Paired Status 1.4B 1.4B 1.4B 160M 160M 160M 1B 1B 1B 2.8B 2.8B 2.8B 410M 410M 410M 70M 70M 70M 0.45 Table 7: Redundancy comparison by Pythia model and seed. Undefined means fewer than two world pairs passed both renderings. 0.30 The table reports log ratios of output size and realized rate, followed by the maximum-budget success difference. Undefined rows remain visible. Paired counts refer to worlds passing in both renderings; success differences use all 200 worlds. System A B C A B C A B C A B C A B C A B C 18 17 21 0 0 0 16 16 17 27 25 26 1 6 3 0 0 0 evaluated evaluated evaluated undefined undefined undefined evaluated evaluated evaluated evaluated evaluated evaluated undefined evaluated evaluated undefined undefined undefined 0.0 0.0 0.0 n/a n/a n/a 0.0 0.0 0.0 0.0 0.0 0.0 n/a 0.0 0.0 n/a n/a n/a ∆ log r ∆ pass Rule -1.229 -1.228 -1.228 n/a n/a n/a -1.230 -1.229 -1.227 -1.231 -1.228 -1.229 n/a -1.224 -1.241 n/a n/a n/a 0.530 0.520 0.520 0.000 0.000 0.000 0.490 0.470 0.490 0.565 0.570 0.555 0.310 0.490 0.505 0.000 0.000 0.000 passes passes passes n/a n/a n/a passes passes passes passes passes passes n/a passes passes n/a n/a n/a The zero output-size log ratios describe the measured size statistic. Equality of output lengths alone leaves serialized content unspecified. D T RANSLATOR SCALE The scale table preserves the individual seed estimates and paired intervals. Pythia 1B and 2.8B have equal byte-axis AUC estimates; their consumer, token-axis, and linter results remain separate measurements. Table 8: Byte-axis closure-AUC by seed and paired difference from Pythia 2.8B. Model 70M 160M A B C Mean ∆ to 2.8B 0.004 0.007 0.003 0.005 -0.667 0.005 0.013 0.006 0.008 -0.663 17 95% interval [-0.678, -0.655] [-0.675, -0.652]
Text of page 18
Continued from previous page Model 410M 1B 1.4B 2.8B Qwen3 0.6B Qwen3 1.7B 0.636 0.672 0.666 0.672 0.587 0.587 0.662 0.672 0.672 0.672 0.586 0.587 C Mean ∆ to 2.8B 0.664 0.672 0.667 0.672 0.586 0.587 Adjacent models 160m minus 70m 410m minus 160m 1b minus 410m 1.4b minus 1b 2.8b minus 1.4b 0.654 0.672 0.668 0.672 0.586 0.587 95% interval -0.018 same AUC -0.003 reference sentinel sentinel [-0.021, -0.015] same AUC [-0.005, -0.002] n/a n/a n/a Difference 95% interval 0.003 0.646 0.018 -0.003 0.003 [0.003, 0.004] [0.635, 0.656] [0.015, 0.021] [-0.005, -0.002] [0.002, 0.005] No model meets the combined AUC and linter equivalence rule. The 410M AUC interval extends past the negative margin. The 1B and 1.4B linter-rate differences exceed the permitted magnitude. Table 10: Full-linter pass-rate differences from the 2.8B reference on the byte axis. B Table 9: Paired AUC differences for adjacent Pythia sizes, larger minus smaller. A Pythia model Linter pass-rate difference 70m 160m 410m 1b 1.4b E -0.029 0.029 -0.084 -0.049 -0.042 C ONSUMER RECOVERY AND STOP REASONS Each consumer receives a code and all canonical entity identifiers, then returns an exact JSON value map. Parser failure scores every entity wrong. The primary output cap is 256 tokens. The secondary GPT-OSS condition changes reasoning effort to low, retaining the prompt and output cap. Recovery is the mean of recorded per-output exact-value fractions across budgets, axes, renderings, and available seeds. The counts include repeated renderings and budget cells, rather than independent worlds. End-turn and token-limit counts expose completion behavior directly. Table 11: Qwen3 Next 80B A3B: exact recovery, parser acceptance, and stop counts. System TE, 1.4B TE, 160M TE, 1B TE, 2.8B Recovery Parser Calls End turn At cap 0.555 0.036 0.586 0.594 18 0.952 0.922 0.964 0.965 5760 5760 5760 5760 5760 5760 5760 5760 0 0 0 0
Text of page 19
Continued from previous page System TE, 410M TE, 70M TE, Qwen3 0.6B TE, Qwen3 1.7B Canonical JSON CCL-Core CCL-Min LLMLingua-2 Oracle core SemanticZip-style Source text Strong SemanticZip-style Structured prose Recovery Parser Calls End turn At cap 0.546 0.025 0.555 0.563 0.703 0.639 0.471 0.169 0.907 0.463 0.869 0.787 0.650 0.952 0.947 0.936 0.935 0.981 0.953 0.966 0.995 1.000 0.968 0.993 1.000 0.952 5760 5760 5760 5760 5760 5760 5760 1920 1920 5760 1920 1920 5760 5760 5760 5760 5760 5760 5760 5760 1920 1920 5760 1920 1920 5760 0 0 0 0 0 0 0 0 0 0 0 0 0 Table 12: GPT-OSS 20B, primary: exact recovery, parser acceptance, and stop counts. System TE, 1.4B TE, 160M TE, 1B TE, 2.8B TE, 410M TE, 70M TE, Qwen3 0.6B TE, Qwen3 1.7B Canonical JSON CCL-Core CCL-Min LLMLingua-2 Oracle core SemanticZip-style Source text Strong SemanticZip-style Structured prose Recovery Parser Calls End turn At cap 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.449 0.000 0.000 0.001 0.001 0.001 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.576 0.000 5760 5760 5760 5760 5760 5760 5760 5760 5760 5760 5760 1920 1920 5760 1920 1920 5760 1 4 3 3 2 0 0 0 0 0 0 0 0 0 0 1120 0 5759 5756 5757 5757 5758 5760 5760 5760 5760 5760 5760 1920 1920 5760 1920 800 5760 Table 13: GPT-OSS 20B, secondary decode control: exact recovery, parser acceptance, and stop counts. System TE, 1.4B TE, 160M TE, 1B TE, 2.8B TE, 410M TE, 70M Recovery Parser Calls End turn At cap 0.097 0.005 0.092 0.109 0.072 0.004 19 0.126 0.149 0.116 0.135 0.094 0.194 5760 5760 5760 5760 5760 5760 2269 2323 2009 2300 1951 2348 3491 3437 3751 3460 3809 3412
Text of page 20
Continued from previous page System TE, Qwen3 0.6B TE, Qwen3 1.7B Canonical JSON CCL-Core CCL-Min LLMLingua-2 Oracle core SemanticZip-style Source text Strong SemanticZip-style Structured prose Recovery Parser Calls End turn At cap 0.083 0.089 0.125 0.095 0.043 0.008 0.206 0.021 0.072 0.743 0.134 0.106 0.116 0.159 0.143 0.056 0.156 0.235 0.032 0.073 0.944 0.196 5760 5760 5760 5760 5760 1920 1920 5760 1920 1920 5760 2333 2503 1862 1679 651 334 452 566 141 1865 1975 3427 3257 3898 4081 5109 1586 1468 5194 1779 55 3785 The primary GPT-OSS condition has extensive truncation and an exception for the strong encoder. The secondary condition also retains token-limit stops. These completion patterns accompany the recovery scores in each condition. Values displayed as 0.000 retain the supplied three-decimal precision and can include nonzero recovery. F E RROR CENSUS The census selects up to 40 error records per series, sorting by the hash of the item identifier. A single item can occur in multiple records through budgets or seeds. An error is a matched-budget closure score below one. Omission and unsupported-fact labels use the penalized explicit precision and recall fields. Budget violations can therefore trigger these labels without identifying the underlying translation error. The linter-valid error label also uses this matched-budget score, so it includes accepted outputs penalized for exceeding their budget. For samples containing budget overruns, marginal counts leave their overlap with accepted outputs unresolved. Five categories require manual inspection and remain unassessed: entity substitution, sign or direction, coefficient or offset, redundancy copying, and unsupported TE fragments. Task-only systems have no closure census. Table 14: Basic score-derived counts. Labels are multi-label counts over the selected records. System Errors Sampled Omission TE, 1.4B TE, 160M TE, 1B TE, 2.8B TE, 410M TE, 70M TE, Qwen3 0.6B TE, Qwen3 1.7B Canonical JSON CCL-Core CCL-Min Oracle core Structured prose 16253 19200 16224 16224 16793 19200 16814 16806 12243 13064 11270 3416 15411 40 40 40 40 40 40 40 40 40 40 40 40 40 20 40 40 40 40 40 40 40 40 40 40 40 40 40 Unsupported Over fact budget 14 40 9 9 16 40 3 4 8 0 19 40 2 5 10 0 0 7 7 0 1 3 0 19 39 2
Text of page 21
Table 15: Parsing and validity counts. Labels are multi-label counts over the selected records. TE, 1.4B TE, 160M TE, 1B TE, 2.8B TE, 410M TE, 70M TE, Qwen3 0.6B TE, Qwen3 1.7B Canonical JSON CCL-Core CCL-Min Oracle core Structured prose G 0 0 0 0 0 2 0 0 0 0 0 0 0 28 35 29 26 32 36 24 25 0 0 0 0 0 12 5 11 14 8 4 16 15 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 12 0 M ACHINE CHECKS OF THE CONSTRUCTED REFERENCE Table 16: Machine validation of the reference comparison. 0 0 0 0 0 2 0 0 5 0 0 0 0 The machine checks compare expected fixture results and serialization stability. Valid Decode Malformed Inconsistent Invalid error fail System H Criterion Evidence Status Machine fixtures Canonical serialization byte-stable (two independent runs of the pilot) Identical closures across the four redundancy variants (six pilot worlds) Linter-valid mistranslation fails the closure comparator Omission and hallucination move the source-conditioned metrics 11/11 identical PASS PASS 6/6 worlds PASS closure agreement 0.167 PASS fixtures 3, 4, 6 PASS A DDITIONAL CONTROL CONSTRUCTION Teacher supervision required a budget-compliant Nova Pro summary whose exact values both consumers recovered. Only 282 of 24,000 assignments passed, leaving too little supervision to train the matched unconstrained-prose control; the three planned cells were removed before grid training. Failures included arithmetic errors, budget overruns, consumer recovery errors and truncation at the 256-token cap, identifier leaks, and transport failures. The frozen records retain the construction results and a separate 24-context natural-text task evaluation, which supplies no closure measurement for the affine comparison. I P ROTOCOL DECISIONS AFFECTING THE COMPARISONS The configuration was fixed using development data before selection or held-out reporting (Nosek et al., 2018). Rank 16 and learning rate 0.0002 were selected in the development sweep. The accelerator choices and removal of the unsuccessful prose control preceded grid training. The secondary GPT-OSS decoding condition was adopted after checkpoint selection and before the report set was opened. The original chronological record and all analysis outputs remain in the reproducibility archive. 21
Text of page 22
J The TE grammar, structural checker, full checker, and training-harness patterns are reused infrastructure. This study adds source-conditioned closure agreement, redundancy response, and translatorscale measurements; related benchmark and consumer-adaptation results contribute no observations to these tables. The independent closure comparator uses decoded equations and exact rational arithmetic to compare source content, including in the linter-valid mistranslation fixture. R EUSED INFRASTRUCTURE 22