---
type: paper
slug: what-survives-learned-symbolic-compression
title: What Survives Learned Symbolic Compression?
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Dmitri Kalaev
- Yanxin Zhang
- Alexey Shvets
date: '2026-08-25'
status: ICLR 2027 submission
line: Telegraph English
pages: 22
html: https://telegrapher.ai/research/what-survives-learned-symbolic-compression/
pdf: https://telegrapher.ai/papers/what-survives-learned-symbolic-compression/what-survives-learned-symbolic-compression.pdf
reader: https://telegrapher.ai/research/what-survives-learned-symbolic-compression/read/
json: https://telegrapher.ai/api/papers/what-survives-learned-symbolic-compression.json
openreview: https://openreview.net/forum?id=hEDqFXFy0v
---

# What Survives Learned Symbolic Compression?

## Paper gist

- **Claim:** A learned symbolic code can pass its own checker and still change what the source said
- **TL;DR:** A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.
- **Method:** On constructed affine-arithmetic worlds (24,000 for training, 200 held out), trace translators trained with low-rank adapters on Pythia checkpoints from 70M to 2.8B parameters (three seeds each) and on two Qwen3 models were compared with Pythia 1.4B encoders trained to write canonical JSON, CCL-Core, CCL-Min and fixed-template prose, and with an irredundant-core oracle, at four byte and four token budgets, by closure agreement with the constructed reference, checker acceptance, and exact value recovery by Qwen3-Next-80B and GPT-OSS-20B.
- **Key result:** Byte-axis closure-AUC, best trace translator: 0.672; Closure agreement of a checker-valid mistranslation: 0.167; Checker-valid outputs among 40 sampled failures: 8–16
- **Why it matters:** Anyone who lets a compressed code stand in for its source has a cheap test to hand: does the code parse and pass its checks?
- **Limits:** The evidence comes from controlled affine arithmetic, whose reference language has no negation, modality, quantifiers, time or uncertainty.
- **Status:** ICLR 2027 submission, August 2026
- **Read:** reader /research/what-survives-learned-symbolic-compression/read/, PDF /papers/what-survives-learned-symbolic-compression/what-survives-learned-symbolic-compression.pdf

## Abstract

Lossy text compression for language-model pipelines is judged by reconstruction distance or by downstream task accuracy, and neither says whether the compressed representation still states what the source stated. We measure agreement with a constructed reference for grammar-constrained symbolic codes. The reference represents source facts as equations, and a fixed decoder recovers the equations a code encodes. An exact rule system compares their consequences independently of the code’s well-formedness checker. On controlled arithmetic micro-worlds, we evaluate learned translators from 70M to 2.8B parameters against structured controls and learned baselines at matched byte and token budgets. Passing the checker is not preserving the source: a self-consistent mistranslation passes it, and so do sampled learned-translation errors. Across byte budgets, the best learned trace translator trails the strongest compact structured control, with area under the closure-agreement curve of 0.67 against 0.86. Increasing translator size does not close this aggregate gap; the largest Pythia translator matches the best smaller translator’s saved byte-axis agreement scores. Source fidelity relative to the constructed reference, internal validity, and what a consumer model recovers are three different quantities and should be measured separately.

## What we did and found

Suppose a source says that x is 2, y is x plus 3, and z is 2 times y. A code that writes y = x + 4 can compute z = 12, check that answer against its own equations, and pass; the source entails z = 10. Nothing inside the code can catch this. For controlled affine-arithmetic worlds, the comparison with the source can be made exact. The source facts become canonical equations, fixed before any code exists. A fixed decoder reads each code format into the same equation language, and exact rational rules expand both sides into a finite set of consequences: the explicit equations, the uniquely entailed values, and the uniquely entailed pairwise differences. The Jaccard overlap of the two sets is closure agreement. In the example the reference has seven consequences, the mistranslation shares two of the twelve in the union, and agreement drops to 0.167. Each code is scored at output allowances of 0.30, 0.45, 0.60 and 0.80 of the nonredundant source, counted once in bytes and once in tokens; an over-budget code scores zero, and the normalised area under the curve is closure-AUC.

Learned translators produce checker-valid source errors as well. Among 40 sampled failing outputs from each of six translators (Pythia 410M to 2.8B and both Qwen3 models), 8 to 16 passed the checker. For Pythia 1B, Pythia 2.8B and Qwen3 0.6B none of the sampled failures was over budget, so the outputs the checker accepted there disagree with the source. Across byte budgets the best trace translator reached a closure-AUC of 0.672. CCL-Min, a compact structured format written by a Pythia 1.4B encoder, reached 0.860: the trace code trailed it under the tighter allowances and caught up near the largest. Size helped, then stopped helping. The 70M and 160M translators stayed near zero, 410M reached 0.654, and 1B, 1.4B and 2.8B sat between 0.668 and 0.672. Counting tokens instead of bytes reversed the order of canonical JSON, CCL-Core and CCL-Min, with canonical JSON ahead at 0.514 and CCL-Min behind at 0.426. Recovery by a consumer model is a separate measurement, taken on a 60-world subset: Qwen3-Next-80B recovered more values from canonical JSON and from the 1B trace code than from CCL-Min, and more from the oracle's explicit equations than from the source text.

## Key numbers

| Measure | Value |
|---|---|
| Byte-axis closure-AUC, best trace translator | 0.672 |
| Closure agreement of a checker-valid mistranslation | 0.167 |
| Checker-valid outputs among 40 sampled failures | 8–16 |
| Token-axis closure-AUC, canonical JSON | 0.514 |
| Exact value recovery by Qwen3-Next-80B, canonical JSON | 0.703 |

## Why it matters

Anyone who lets a compressed code stand in for its source has a cheap test to hand: does the code parse and pass its checks? The fixture and the sampled learned errors show what that test misses. A checker compares the code with itself, and a consistent mistranslation is, after all, consistent. Evidence that the content survived needs a reference on the source side, fixed before any code is written, plus a decoder that puts the code into the same terms.

Which format wins depends on what is counted. CCL-Min comes out ahead per byte and canonical JSON per token, and the oracle, holding the correct facts by construction, still scores zero at the tightest byte allowance because its complete code does not fit. Past 1B, more translator capacity bought nothing on the byte axis, while a different target format did: a 1.4B encoder writing CCL-Min retained more per byte than any trace translator, the 2.8B included. An encoding comparison tells a pipeline something when its budget unit is the resource that pipeline actually spends.

The trace format reuses the Telegraph English grammar and checker, so the question lands on that line of work. Telegraph English and Context Compression Is Not One Thing judged symbolic re-expression by whether a model could still answer questions over it; this paper adds the source-side check that question answering leaves out. Its instrument has a relative in RippleKB, which scores systems against a constructed dependency closure of the source units one edited fact affects.

## What this does not show

The evidence comes from controlled affine arithmetic, whose reference language has no negation, modality, quantifiers, time or uncertainty. Moving to open text needs a source representation and a consequence family chosen for that domain, and the finite signature already builds such choices into the score: logically equivalent equation sets can score differently. The checker-failure samples are selected on failure, so they show that accepted source errors occur, not how often, and the error categories that need manual inspection, offset changes like the fixture's among them, were left unassessed. The size plateau holds for one training and selection procedure, and the byte-axis curves of the 1B and 2.8B translators coincide exactly while their token-axis and checker measurements differ. Consumer recovery was measured on a 60-world subset across both budget axes, so its ordering cannot be set against closure agreement on the same codes; GPT-OSS-20B's recovery rounded to zero for every input but one under the 256-token response cap and rose once its reasoning effort was lowered, which makes recovery a function of decoding settings too. A planned unconstrained-prose control was dropped before training because 282 of 24,000 teacher assignments passed, too few to supervise it. The release holds aggregate tables, configuration and checkpoint hashes and the analysis code, but not raw outputs, weights, or the imported construction and decoding code, so the adapted CCL serialization and raw-output equality cannot be checked independently.

## Blog post

[A compressed code can pass its own checker and still change the facts](https://telegrapher.ai/blog/what-survives-learned-symbolic-compression.md)

## Cite

```bibtex
@misc{bei2026what,
  title         = {What Survives Learned Symbolic Compression?},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/what-survives-learned-symbolic-compression/}
}
```
