telegrapher

Research for people who ship language models.

Error-accumulation

All three papers

How errors build up as a model generates, where they cluster, and how production systems correct for them outside the weights. Read in order: the three papers form one argument.

Telegraph English

5 papers

Rewriting prompts and documents into a compact symbolic form, and measuring which facts and relations survive the rewrite.

What Survives Learned Symbolic Compression?

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.

ICLR 2027 submissionAugust 2026Blog post

Evaluating Relational Context Compression at Realized Token Budgets

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths.

ICLR 2027 submissionSeptember 2026Blog post

RippleKB: Finding What an Edit Changes Across Linked Documents

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev

Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.

ICLR 2027 submissionAugust 2026Blog post

People

Everyone

Verifiable reasoning

3 papers

Reasoning traces that can be checked mechanically, and the measurements needed to tell whether training actually improves them.

Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.

ICLR 2027 submissionAugust 2026Blog post

Conversation analytics

4 papers

Labelling, categorising and normalising customer conversations with LLMs under real budgets, and auditing the labels without a gold set.

Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text

Navita Jain, Mikhail L Arbuzov, Dmitry Dimov, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Karan Dave, Sisong Bei

Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not.

ICLR 2027 submissionSeptember 2026Blog post

Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale

Mikhail L Arbuzov, Sisong Bei, Dmitry Dimov, Karan Dave, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Navita Jain

Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path.

ICLR 2027 submissionSeptember 2026Blog post

From the blog

All posts

Long LLM outputs hinge on a few decisions, not on their length

About 9% of tokens carry long-range dependency and errors cluster, so the predicted decay of long LLM outputs is far gentler than (1 − e)^n.

October 9, 2026. On Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

No finite fix list covers every LLM failure, and a deployment doesn't need one

Open-ended LLM use has no finite fix list, but inside one deployment failures tend to recur, so a slowly growing library of fixes covers the most common ones.

October 9, 2026. On The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability

Production AI learns outside the weights, without the optimizer that layer needs

Teams fix LLM systems by editing prompts, memory, skills and tools. Across ~130 systems, that loop runs mostly as patchwork, its cross-org optimizer missing.

October 9, 2026. On Frontier and Localhost: How Production AI Learns Outside the Weights

Rewriting a prompt loses fewer details than deleting its tokens

Telegraph English rewrites text one claim per line with explicit symbols. Against LLMLingua-2 it lost fewer answers, most visibly on fine-detail questions.

October 9, 2026. On Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting