# Ziwei Dong

Error accumulation, compression and verifiable reasoning.

Ziwei Dong is a co-author across the error-accumulation trilogy, the Telegraph English and compression papers, and the reasoning papers.

## Papers

- [Evaluating Relational Context Compression at Realized Token Budgets](https://telegrapher.ai/research/relational-context-compression.md): Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths. (ICLR 2027 submission, September 2026)
- [What Survives Learned Symbolic Compression?](https://telegrapher.ai/research/what-survives-learned-symbolic-compression.md): A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open. (ICLR 2027 submission, August 2026)
- [RippleKB: Finding What an Edit Changes Across Linked Documents](https://telegrapher.ai/research/ripplekb.md): Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself. (ICLR 2027 submission, August 2026)
- [Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning](https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability.md): On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy. (ICLR 2027 submission, August 2026)
- [The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability](https://telegrapher.ai/research/architecture-of-errors.md): No finite fix list covers every failure mode of open-ended LLM use. Inside one deployment, published taxonomies suggest failures recur in a small catalogue, so a sufficient fix library grows slowly with sequence length, then levels off. (COLM 2026 workshop poster, August 2026)
- [Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It](https://telegrapher.ai/research/additive-process-rewards.md): Under GRPO's per-group standardization, a process term added to the outcome reward cannot be tuned: its weight cancels where a group's outcomes agree and is swamped where they differ. Endpoint gains need replication and a shuffle control. (AAAI 2027 submission, July 2026)
- [Frontier and Localhost: How Production AI Learns Outside the Weights](https://telegrapher.ai/research/frontier-and-localhost.md): Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations. (COLM 2026 workshop submission, June 2026)
- [Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget](https://telegrapher.ai/research/context-compression-is-not-one-thing.md): Rewriting retrieved passages as pipe-separated entity-relation clauses, entities kept verbatim, beat character deletion, truncation and random subsampling at the same token budget on three multi-hop benchmarks. At a fixed budget, the form of the kept text is not a detail. (ACL ARR 2026 submission, May 2026)
- [Telegraph Reasoning: Lintable Traces for Mechanically Verified Chain-of-Thought](https://telegrapher.ai/research/telegraph-reasoning.md): If a model writes its math reasoning as tagged lines, a rule-based linter can recompute equations and check claims in SymPy instead of asking a model to reread them. It caught 195 of 196 injected errors; self-verification, about two thirds. (Preprint, May 2026)
- [Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models](https://telegrapher.ai/research/beyond-exponential-decay.md): Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length. (NeurIPS 2026 submission, May 2026)
- [Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting](https://telegrapher.ai/research/telegraph-english.md): Rewriting text so each line holds one claim, with explicit symbols for cause and contrast, cut tokens by about two fifths and lost fewer answers than LLMLingua-2's token deletion, by the widest margin on fine details. (NeurIPS 2026 submission, May 2026)
