# Telegrapher > We study where language models actually break, and how to fix them without retraining. Telegrapher is an independent, unfunded, unincorporated research collaboration. Research for people who ship language models. How long outputs actually break. How to shrink a prompt without losing the facts. How a deployed system keeps improving without retraining. Every paper comes with a plain-language write-up, and a version your agents can read. Every page on telegrapher.ai has a markdown twin: replace the trailing slash with `.md` (for example https://telegrapher.ai/research/telegraph-english.md). JSON: https://telegrapher.ai/api/index.json lists everything; each paper is also at `/api/papers/.json` with its gist. PDFs are at `/papers//.pdf`. ## Error-accumulation How errors build up as a model generates, where they cluster, and how production systems correct for them outside the weights. Read in order: the three papers form one argument. - [Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models](https://telegrapher.ai/research/beyond-exponential-decay.md): Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length. (NeurIPS 2026 submission, May 2026) - [The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability](https://telegrapher.ai/research/architecture-of-errors.md): No finite fix list covers every failure mode of open-ended LLM use. Inside one deployment, published taxonomies suggest failures recur in a small catalogue, so a sufficient fix library grows slowly with sequence length, then levels off. (COLM 2026 workshop poster, August 2026) - [Frontier and Localhost: How Production AI Learns Outside the Weights](https://telegrapher.ai/research/frontier-and-localhost.md): Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations. (COLM 2026 workshop submission, June 2026) ## Telegraph English Rewriting prompts and documents into a compact symbolic form, and measuring which facts and relations survive the rewrite. - [Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting](https://telegrapher.ai/research/telegraph-english.md): Rewriting text so each line holds one claim, with explicit symbols for cause and contrast, cut tokens by about two fifths and lost fewer answers than LLMLingua-2's token deletion, by the widest margin on fine details. (NeurIPS 2026 submission, May 2026) - [Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget](https://telegrapher.ai/research/context-compression-is-not-one-thing.md): Rewriting retrieved passages as pipe-separated entity-relation clauses, entities kept verbatim, beat character deletion, truncation and random subsampling at the same token budget on three multi-hop benchmarks. At a fixed budget, the form of the kept text is not a detail. (ACL ARR 2026 submission, May 2026) - [What Survives Learned Symbolic Compression?](https://telegrapher.ai/research/what-survives-learned-symbolic-compression.md): A symbolic code's own checker accepted translations that changed the source. Scored against a constructed reference, trace translators kept less per byte than a compact structured format, and larger translators left the gap open. (ICLR 2027 submission, August 2026) - [Evaluating Relational Context Compression at Realized Token Budgets](https://telegrapher.ai/research/relational-context-compression.md): Asked to rewrite passages into relational text at a set token budget, two LLM encoders landed in the registered band on 33 of 9,600 outputs, and their median output ran long. Compressors should be compared at delivered lengths. (ICLR 2027 submission, September 2026) - [RippleKB: Finding What an Edit Changes Across Linked Documents](https://telegrapher.ai/research/ripplekb.md): Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself. (ICLR 2027 submission, August 2026) ## Verifiable reasoning Reasoning traces that can be checked mechanically, and the measurements needed to tell whether training actually improves them. - [Telegraph Reasoning: Lintable Traces for Mechanically Verified Chain-of-Thought](https://telegrapher.ai/research/telegraph-reasoning.md): If a model writes its math reasoning as tagged lines, a rule-based linter can recompute equations and check claims in SymPy instead of asking a model to reread them. It caught 195 of 196 injected errors; self-verification, about two thirds. (Preprint, May 2026) - [Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning](https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability.md): On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy. (ICLR 2027 submission, August 2026) - [Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning — and What It Takes to Measure It](https://telegrapher.ai/research/additive-process-rewards.md): Under GRPO's per-group standardization, a process term added to the outcome reward cannot be tuned: its weight cancels where a group's outcomes agree and is swamped where they differ. Endpoint gains need replication and a shuffle control. (AAAI 2027 submission, July 2026) ## Conversation analytics Labelling, categorising and normalising customer conversations with LLMs under real budgets, and auditing the labels without a gold set. - [The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling](https://telegrapher.ai/research/same-family-halo.md): An LLM labeler agrees more with silver labels written by a sibling model than with labels from other families. The contrast needs no human gold, so same-family consensus can be audited, and discounted, where gold is scarce. (ICLR 2027 submission, September 2026) - [High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution](https://telegrapher.ai/research/high-load-call-categorization.md): A confidence gate lets an encoder answer roughly 87% of customer calls; with a five-label shortlist for the rest, cost falls more than 90% below an LLM reading the full list on every call. The shortlist needs its own prompt. (ICLR 2027 submission, September 2026) - [Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text](https://telegrapher.ai/research/clusters-are-proposals.md): Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not. (ICLR 2027 submission, September 2026) - [Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale](https://telegrapher.ai/research/clarify-then-focus.md): Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path. (ICLR 2027 submission, September 2026) ## People - [Sisong Bei](https://telegrapher.ai/people/sisong-bei.md): Compression, verifiable reasoning and what a model can still do with less. - [Mikhail L Arbuzov](https://telegrapher.ai/people/mikhail-arbuzov.md): How errors accumulate in LLM systems, and how much of a prompt a model needs. - [Ziwei Dong](https://telegrapher.ai/people/ziwei-dong.md): Error accumulation, compression and verifiable reasoning. - [Dmitri Kalaev](https://telegrapher.ai/people/dmitri-kalaev.md): Error accumulation, compression and verifiable reasoning. - [Lee Mosbacker](https://telegrapher.ai/people/lee-mosbacker.md): Telegraph English and how production AI systems learn outside the weights. - [Alexey Shvets](https://telegrapher.ai/people/alexey-shvets.md): Error accumulation, compression and verifiable reasoning. ## Blog - [Long LLM outputs hinge on a few decisions, not on their length](https://telegrapher.ai/blog/beyond-exponential-decay.md): About 9% of tokens carry long-range dependency and errors cluster, so the predicted decay of long LLM outputs is far gentler than (1 − e)^n. - [No finite fix list covers every LLM failure, and a deployment doesn't need one](https://telegrapher.ai/blog/architecture-of-errors.md): Open-ended LLM use has no finite fix list, but inside one deployment failures tend to recur, so a slowly growing library of fixes covers the most common ones. - [Production AI learns outside the weights, without the optimizer that layer needs](https://telegrapher.ai/blog/frontier-and-localhost.md): Teams fix LLM systems by editing prompts, memory, skills and tools. Across ~130 systems, that loop runs mostly as patchwork, its cross-org optimizer missing. - [Rewriting a prompt loses fewer details than deleting its tokens](https://telegrapher.ai/blog/telegraph-english.md): Telegraph English rewrites text one claim per line with explicit symbols. Against LLMLingua-2 it lost fewer answers, most visibly on fine-detail questions. - [Rewriting a retrieved passage beats cutting it to the same token budget](https://telegrapher.ai/blog/context-compression-is-not-one-thing.md): At a fixed token budget, rewriting passages as entity-relation clauses beat truncated and thinned text on three multi-hop datasets, and a summary on one. - [A compressed code can pass its own checker and still change the facts](https://telegrapher.ai/blog/what-survives-learned-symbolic-compression.md): Learned symbolic codes passed their own checker while changing what the source said. Scored against the source, a compact format kept more per byte. - [Asked for a quarter of the context, one encoder sent back two thirds](https://telegrapher.ai/blog/relational-context-compression.md): We followed Telegraph English rewrites from token budget to reader: 33 of 9,600 landed in band, and a compiler that held its ceiling still scored poorly. - [Finding most of what an edit changed is not finding all of it](https://telegrapher.ai/blog/ripplekb.md): We scored whether rankings recover every statement an edit changes, not just most. Recall and completion disagree, sometimes at the edited sentence itself. - [A linter that redoes the work catches injected errors that self-verification misses](https://telegrapher.ai/blog/telegraph-reasoning.md): Written in seven tags, a model's math reasoning can be checked by a SymPy-backed linter. It caught 195 of 196 injected errors; self-verification, 66.3%. - [A math trace that passes our checker is right about one time in three](https://telegrapher.ai/blog/answer-accuracy-and-trace-verifiability.md): We checked a small model's math traces with a program. Accepted traces were right about one time in three; switching to the trace format cost accuracy. - [In GRPO, the weight on an additive process reward acts like a switch, not a dial](https://telegrapher.ai/blog/additive-process-rewards.md): Group standardization cancels an additive process reward's weight when outcomes agree and swamps the term when they differ. We measured both, then tested fixes. - [An LLM labeler scores higher when its sibling wrote the answer key](https://telegrapher.ai/blog/same-family-halo.md): Hold an LLM labeler fixed and swap who wrote its answer key: agreement rises when the key comes from a sibling model, and no human gold is needed to see it. - [Most calls never reach the LLM, and the ones that do need a different prompt](https://telegrapher.ai/blog/high-load-call-categorization.md): A fine-tuned encoder answers roughly 87% of customer calls. The rest go to an LLM with five candidate labels, and that shortlist needs a prompt of its own. - [Rare call issues can get lost before an LLM ever sees them](https://telegrapher.ai/blog/clusters-are-proposals.md): A 1% call issue can be absent from the sample an LLM proposes topics from. Over-splitting each category found 234 subcategories; flat clustering found 111. - [Clarify a call once, then choose what each reader sees](https://telegrapher.ai/blog/clarify-then-focus.md): We rewrite each customer call once into short, tagged statements. An encoder gains from clearer text alone; weaker prompted readers also gain from selection. ## Optional - [About](https://telegrapher.ai/about.md): who we are and how the group works