telegrapher

Alexey Shvets

Error accumulation, compression and verifiable reasoning.

11papers 3research lines

About

Alexey Shvets is a co-author across the error-accumulation trilogy, the Telegraph English and compression papers, and the reasoning papers.

to fill: Affiliation (confirm wording), bio, photo, links.

Papers

Evaluating Relational Context Compression at Realized Token Budgets

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths.

ICLR 2027 submissionSeptember 2026Telegraph EnglishBlog post

What Survives Learned Symbolic Compression?

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.

ICLR 2027 submissionAugust 2026Telegraph EnglishBlog post

RippleKB: Finding What an Edit Changes Across Linked Documents

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev

Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.

ICLR 2027 submissionAugust 2026Telegraph EnglishBlog post

Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.

ICLR 2027 submissionAugust 2026Verifiable reasoningBlog post

Frontier and Localhost: How Production AI Learns Outside the Weights

Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets

Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations.

COLM 2026 workshop submissionJune 2026Error-accumulationBlog post

Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev

Rewriting retrieved passages as pipe-separated entity-relation clauses, entities kept verbatim, beat character deletion, truncation and random subsampling at the same token budget on three multi-hop benchmarks. At a fixed budget, the form of the kept text is not a detail.

ACL ARR 2026 submissionMay 2026Telegraph EnglishBlog post

Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets

Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length.

NeurIPS 2026 submissionMay 2026Error-accumulationBlog post

Blog posts

Long LLM outputs hinge on a few decisions, not on their length

About 9% of tokens carry long-range dependency and errors cluster, so the predicted decay of long LLM outputs is far gentler than (1 − e)^n.

October 9, 2026. On Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

No finite fix list covers every LLM failure, and a deployment doesn't need one

Open-ended LLM use has no finite fix list, but inside one deployment failures tend to recur, so a slowly growing library of fixes covers the most common ones.

October 9, 2026. On The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability

Production AI learns outside the weights, without the optimizer that layer needs

Teams fix LLM systems by editing prompts, memory, skills and tools. Across ~130 systems, that loop runs mostly as patchwork, its cross-org optimizer missing.

October 9, 2026. On Frontier and Localhost: How Production AI Learns Outside the Weights

Rewriting a prompt loses fewer details than deleting its tokens

Telegraph English rewrites text one claim per line with explicit symbols. Against LLMLingua-2 it lost fewer answers, most visibly on fine-detail questions.

October 9, 2026. On Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting

Rewriting a retrieved passage beats cutting it to the same token budget

At a fixed token budget, rewriting passages as entity-relation clauses beat truncated and thinned text on three multi-hop datasets, and a summary on one.

October 9, 2026. On Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

A compressed code can pass its own checker and still change the facts

Learned symbolic codes passed their own checker while changing what the source said. Scored against the source, a compact format kept more per byte.

October 9, 2026. On What Survives Learned Symbolic Compression?

Asked for a quarter of the context, one encoder sent back two thirds

We followed Telegraph English rewrites from token budget to reader: 33 of 9,600 landed in band, and a compiler that held its ceiling still scored poorly.

October 9, 2026. On Evaluating Relational Context Compression at Realized Token Budgets

Finding most of what an edit changed is not finding all of it

We scored whether rankings recover every statement an edit changes, not just most. Recall and completion disagree, sometimes at the edited sentence itself.

October 9, 2026. On RippleKB: Finding What an Edit Changes Across Linked Documents

A linter that redoes the work catches injected errors that self-verification misses

Written in seven tags, a model's math reasoning can be checked by a SymPy-backed linter. It caught 195 of 196 injected errors; self-verification, 66.3%.

October 9, 2026. On Telegraph Reasoning: Lintable Traces for Mechanically Verified Chain-of-Thought

A math trace that passes our checker is right about one time in three

We checked a small model's math traces with a program. Accepted traces were right about one time in three; switching to the trace format cost accuracy.

October 9, 2026. On Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

In GRPO, the weight on an additive process reward acts like a switch, not a dial

Group standardization cancels an additive process reward's weight when outcomes agree and swamps the term when they differ. We measured both, then tested fixes.

October 9, 2026. On Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It