telegrapher

Research

15 papers in four lines of work. Each paper page has a short summary, a page-by-page reader and the PDF.

No paper matches that filter. Clear the box or pick another line.

Error-accumulation

How errors build up as a model generates, where they cluster, and how production systems correct for them outside the weights. Read in order: the three papers form one argument.

Frontier and Localhost: How Production AI Learns Outside the Weights

Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets

Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations.

COLM 2026 workshop submissionJune 2026Blog post

Telegraph English

Rewriting prompts and documents into a compact symbolic form, and measuring which facts and relations survive the rewrite.

What Survives Learned Symbolic Compression?

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open.

ICLR 2027 submissionAugust 2026Blog post

Evaluating Relational Context Compression at Realized Token Budgets

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths.

ICLR 2027 submissionSeptember 2026Blog post

RippleKB: Finding What an Edit Changes Across Linked Documents

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev

Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.

ICLR 2027 submissionAugust 2026Blog post

Verifiable reasoning

Reasoning traces that can be checked mechanically, and the measurements needed to tell whether training actually improves them.

Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning

Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Yan Han, Dmitri Kalaev, Yanxin Zhang, Alexey Shvets

On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.

ICLR 2027 submissionAugust 2026Blog post

Conversation analytics

Labelling, categorising and normalising customer conversations with LLMs under real budgets, and auditing the labels without a gold set.

Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text

Navita Jain, Mikhail L Arbuzov, Dmitry Dimov, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Karan Dave, Sisong Bei

Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not.

ICLR 2027 submissionSeptember 2026Blog post

Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale

Mikhail L Arbuzov, Sisong Bei, Dmitry Dimov, Karan Dave, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Navita Jain

Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path.

ICLR 2027 submissionSeptember 2026Blog post