telegrapher

Mikhail L Arbuzov

Independent researcher

How errors accumulate in LLM systems, and how much of a prompt a model needs.

15papers 7as first author4research lines

About

Mikhail L Arbuzov works on two questions that turned out to be connected: how errors build up as a language model generates, and how much of a prompt the model actually needs to read. The first produced the error-accumulation trilogy, which moves from a critique of the exponential-decay picture, through a theory of where errors cluster, to how production systems learn outside the model weights. The second produced Telegraph English, a structured symbolic rewrite of prompts.

Mikhail is first author on Beyond Exponential Decay, The Architecture of Errors, Frontier and Localhost, Telegraph English, and three applied papers on conversation analytics at scale (The Same-Family Halo, High-Load Budgeted Categorization, and Clarify, Then Focus). Mikhail is also a co-author on the compression and reasoning papers and on Clusters Are Proposals.

to fill: Background, current focus, photo, personal links.

Papers

Evaluating Relational Context Compression at Realized Token Budgets

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths.

ICLR 2027 submissionSeptember 2026Telegraph EnglishBlog post

High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution

Mikhail L Arbuzov, Sisong Bei, Dmitry Dimov, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Karan Dave, Navita Jain

A confidence gate lets an encoder answer roughly 87% of customer calls; with a five-label shortlist for the rest, cost falls more than 90% below an LLM reading the full list on every call. The shortlist needs its own prompt.

ICLR 2027 submissionSeptember 2026Conversation analyticsBlog post

Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text

Navita Jain, Mikhail L Arbuzov, Dmitry Dimov, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Karan Dave, Sisong Bei

Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not.

ICLR 2027 submissionSeptember 2026Conversation analyticsBlog post

Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale

Mikhail L Arbuzov, Sisong Bei, Dmitry Dimov, Karan Dave, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Navita Jain

Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path.

ICLR 2027 submissionSeptember 2026Conversation analyticsBlog post

What Survives Learned Symbolic Compression?

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.

ICLR 2027 submissionAugust 2026Telegraph EnglishBlog post

RippleKB: Finding What an Edit Changes Across Linked Documents

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev

Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.

ICLR 2027 submissionAugust 2026Telegraph EnglishBlog post

Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.

ICLR 2027 submissionAugust 2026Verifiable reasoningBlog post

Frontier and Localhost: How Production AI Learns Outside the Weights

Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets

Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations.

COLM 2026 workshop submissionJune 2026Error-accumulationBlog post

Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev

Rewriting retrieved passages as pipe-separated entity-relation clauses, entities kept verbatim, beat character deletion, truncation and random subsampling at the same token budget on three multi-hop benchmarks. At a fixed budget, the form of the kept text is not a detail.

ACL ARR 2026 submissionMay 2026Telegraph EnglishBlog post

Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets

Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length.

NeurIPS 2026 submissionMay 2026Error-accumulationBlog post

Blog posts

Long LLM outputs hinge on a few decisions, not on their length

About 9% of tokens carry long-range dependency and errors cluster, so the predicted decay of long LLM outputs is far gentler than (1 − e)^n.

October 9, 2026. On Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

No finite fix list covers every LLM failure, and a deployment doesn't need one

Open-ended LLM use has no finite fix list, but inside one deployment failures tend to recur, so a slowly growing library of fixes covers the most common ones.

October 9, 2026. On The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability

Production AI learns outside the weights, without the optimizer that layer needs

Teams fix LLM systems by editing prompts, memory, skills and tools. Across ~130 systems, that loop runs mostly as patchwork, its cross-org optimizer missing.

October 9, 2026. On Frontier and Localhost: How Production AI Learns Outside the Weights

Rewriting a prompt loses fewer details than deleting its tokens

Telegraph English rewrites text one claim per line with explicit symbols. Against LLMLingua-2 it lost fewer answers, most visibly on fine-detail questions.

October 9, 2026. On Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting

Rewriting a retrieved passage beats cutting it to the same token budget

At a fixed token budget, rewriting passages as entity-relation clauses beat truncated and thinned text on three multi-hop datasets, and a summary on one.

October 9, 2026. On Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

A compressed code can pass its own checker and still change the facts

Learned symbolic codes passed their own checker while changing what the source said. Scored against the source, a compact format kept more per byte.

October 9, 2026. On What Survives Learned Symbolic Compression?

Asked for a quarter of the context, one encoder sent back two thirds

We followed Telegraph English rewrites from token budget to reader: 33 of 9,600 landed in band, and a compiler that held its ceiling still scored poorly.

October 9, 2026. On Evaluating Relational Context Compression at Realized Token Budgets

Finding most of what an edit changed is not finding all of it

We scored whether rankings recover every statement an edit changes, not just most. Recall and completion disagree, sometimes at the edited sentence itself.

October 9, 2026. On RippleKB: Finding What an Edit Changes Across Linked Documents

A linter that redoes the work catches injected errors that self-verification misses

Written in seven tags, a model's math reasoning can be checked by a SymPy-backed linter. It caught 195 of 196 injected errors; self-verification, 66.3%.

October 9, 2026. On Telegraph Reasoning: Lintable Traces for Mechanically Verified Chain-of-Thought

A math trace that passes our checker is right about one time in three

We checked a small model's math traces with a program. Accepted traces were right about one time in three; switching to the trace format cost accuracy.

October 9, 2026. On Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

In GRPO, the weight on an additive process reward acts like a switch, not a dial

Group standardization cancels an additive process reward's weight when outcomes agree and swamps the term when they differ. We measured both, then tested fixes.

October 9, 2026. On Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It

An LLM labeler scores higher when its sibling wrote the answer key

Hold an LLM labeler fixed and swap who wrote its answer key: agreement rises when the key comes from a sibling model, and no human gold is needed to see it.

October 9, 2026. On The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling

Most calls never reach the LLM, and the ones that do need a different prompt

A fine-tuned encoder answers roughly 87% of customer calls. The rest go to an LLM with five candidate labels, and that shortlist needs a prompt of its own.

October 9, 2026. On High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution

Rare call issues can get lost before an LLM ever sees them

A 1% call issue can be absent from the sample an LLM proposes topics from. Over-splitting each category found 234 subcategories; flat clustering found 111.

October 9, 2026. On Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text

Clarify a call once, then choose what each reader sees

We rewrite each customer call once into short, tagged statements. An encoder gains from clearer text alone; weaker prompted readers also gain from selection.

October 9, 2026. On Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale