A learned symbolic code can pass its own checker and still change what the source said
What Survives Learned Symbolic Compression?
Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets
ICLR 2027 submission, August 2026
A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.
What we did and found
Suppose a source says that x is 2, y is x plus 3, and z is 2 times y. A code that writes y = x + 4 can compute z = 12, check that answer against its own equations, and pass; the source entails z = 10. Nothing inside the code can catch this. For controlled affine-arithmetic worlds, the comparison with the source can be made exact. The source facts become canonical equations, fixed before any code exists. A fixed decoder reads each code format into the same equation language, and exact rational rules expand both sides into a finite set of consequences: the explicit equations, the uniquely entailed values, and the uniquely entailed pairwise differences. The Jaccard overlap of the two sets is closure agreement. In the example the reference has seven consequences, the mistranslation shares two of the twelve in the union, and agreement drops to 0.167. Each code is scored at output allowances of 0.30, 0.45, 0.60 and 0.80 of the nonredundant source, counted once in bytes and once in tokens; an over-budget code scores zero, and the normalised area under the curve is closure-AUC.
Learned translators produce checker-valid source errors as well. Among 40 sampled failing outputs from each of six translators (Pythia 410M to 2.8B and both Qwen3 models), 8 to 16 passed the checker. For Pythia 1B, Pythia 2.8B and Qwen3 0.6B none of the sampled failures was over budget, so the outputs the checker accepted there disagree with the source. Across byte budgets the best trace translator reached a closure-AUC of 0.672. CCL-Min, a compact structured format written by a Pythia 1.4B encoder, reached 0.860: the trace code trailed it under the tighter allowances and caught up near the largest. Size helped, then stopped helping. The 70M and 160M translators stayed near zero, 410M reached 0.654, and 1B, 1.4B and 2.8B sat between 0.668 and 0.672. Counting tokens instead of bytes reversed the order of canonical JSON, CCL-Core and CCL-Min, with canonical JSON ahead at 0.514 and CCL-Min behind at 0.426. Recovery by a consumer model is a separate measurement, taken on a 60-world subset: Qwen3-Next-80B recovered more values from canonical JSON and from the 1B trace code than from CCL-Min, and more from the oracle's explicit equations than from the source text.
Key numbers
| Byte-axis closure-AUC, best trace translatorPythia 1B and 2.8B; CCL-Min, written by a Pythia 1.4B encoder, reaches 0.860 over the same four byte budgets | 0.672 |
| Closure agreement of a checker-valid mistranslationconstructed fixture with y = x + 4 in place of y = x + 3; the exact code scores 1.000 | 0.167 |
| Checker-valid outputs among 40 sampled failuresper translator, for Pythia 410M to 2.8B and both Qwen3 models; samples are selected on failure, so they show the case exists, not its rate | 8–16 |
| Token-axis closure-AUC, canonical JSONahead of the other structured controls on tokens; CCL-Min falls to 0.426 and trace translators reach at most 0.369 | 0.514 |
| Exact value recovery by Qwen3-Next-80B, canonical JSONagainst 0.471 for CCL-Min and 0.869 for the source text, on the 60-world consumer subset | 0.703 |
What this does not show
The evidence comes from controlled affine arithmetic, whose reference language has no negation, modality, quantifiers, time or uncertainty. Moving to open text needs a source representation and a consequence family chosen for that domain, and the finite signature already builds such choices into the score: logically equivalent equation sets can score differently. The checker-failure samples are selected on failure, so they show that accepted source errors occur, not how often, and the error categories that need manual inspection, offset changes like the fixture's among them, were left unassessed. The size plateau holds for one training and selection procedure, and the byte-axis curves of the 1B and 2.8B translators coincide exactly while their token-axis and checker measurements differ. Consumer recovery was measured on a 60-world subset across both budget axes, so its ordering cannot be set against closure agreement on the same codes; GPT-OSS-20B's recovery rounded to zero for every input but one under the 256-token response cap and rose once its reasoning effort was lowered, which makes recovery a function of decoding settings too. A planned unconstrained-prose control was dropped before training because 282 of 24,000 teacher assignments passed, too few to supervise it. The release holds aggregate tables, configuration and checkpoint hashes and the analysis code, but not raw outputs, weights, or the imported construction and decoding code, so the adapted CCL serialization and raw-output equality cannot be checked independently.
The authors' abstract
Lossy text compression for language-model pipelines is judged by reconstruction distance or by downstream task accuracy, and neither says whether the compressed representation still states what the source stated. We measure agreement with a constructed reference for grammar-constrained symbolic codes. The reference represents source facts as equations, and a fixed decoder recovers the equations a code encodes. An exact rule system compares their consequences independently of the code’s well-formedness checker. On controlled arithmetic micro-worlds, we evaluate learned translators from 70M to 2.8B parameters against structured controls and learned baselines at matched byte and token budgets. Passing the checker is not preserving the source: a self-consistent mistranslation passes it, and so do sampled learned-translation errors. Across byte budgets, the best learned trace translator trails the strongest compact structured control, with area under the closure-agreement curve of 0.67 against 0.86. Increasing translator size does not close this aggregate gap; the largest Pythia translator matches the best smaller translator’s saved byte-axis agreement scores. Source fidelity relative to the constructed reference, internal validity, and what a consumer model recovers are three different quantities and should be measured separately.
More in Telegraph English
Cite
@misc{bei2026what,
title = {What Survives Learned Symbolic Compression?},
author = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},
year = {2026},
note = {ICLR 2027 submission},
url = {https://telegrapher.ai/research/what-survives-learned-symbolic-compression/}
}Builds on
- Xu (2026). Semantic rate-distortion theory: Deductive compression and closure fidelity.
- Trukhina and Vashkelis (2026). Compress the context, keep the commitments: A formal framework for verifiable LLM context compression.
- Trukhina and Vashkelis (2026). SemanticZip: A pilot framework for lossy text compression with LLMs as semantic decompressors.
- Arbuzov et al. (2026). Telegraph English: Semantic prompt compression via structured symbolic rewriting.
- Pan et al. (2024). LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression.