---
type: paper
slug: beyond-exponential-decay
title: 'Beyond Exponential Decay: Rethinking Error Accumulation in Large Language
  Models'
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Ziwei Dong
- Dmitri Kalaev
- Alexey Shvets
date: '2026-05-04'
status: NeurIPS 2026 submission
line: Error-accumulation
pages: 15
html: https://telegrapher.ai/research/beyond-exponential-decay/
pdf: https://telegrapher.ai/papers/beyond-exponential-decay/beyond-exponential-decay.pdf
reader: https://telegrapher.ai/research/beyond-exponential-decay/read/
json: https://telegrapher.ai/api/papers/beyond-exponential-decay.json
arxiv: https://arxiv.org/abs/2505.24187
openreview: https://openreview.net/forum?id=cZl2t7D4m6
---

# Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

## Paper gist

- **Claim:** Long LLM outputs hinge on a few key tokens, so predicted decay is far gentler than (1 − e)^n
- **TL;DR:** Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length.
- **Method:** A synthesis with no new experiments: the paper folds published measurements of key-token sparsity, embedding geometry and ensemble convergence into a two-rate error model, then re-reads existing systems results (anchor compression, RetrievalAttention, self-consistency, tool-integrated reasoning) through it.
- **Key result:** Naive 100-token correctness: 37%; Tokens that need long-range context: 9%; Key-token perplexity vs. task performance: −0.96
- **Why it matters:** For engineers building long-context or agentic systems, the practical change is the unit of the reliability budget: decisions, not tokens.
- **Limits:** The framework is a synthesis, not a derivation: the paper runs no experiments of its own, and its measured figures come from the cited works.
- **Status:** NeurIPS 2026 submission, May 2026
- **Read:** reader /research/beyond-exponential-decay/read/, PDF /papers/beyond-exponential-decay/beyond-exponential-decay.pdf, arXiv:2505.24187 https://arxiv.org/abs/2505.24187

## Abstract

A common pessimistic argument holds that autoregressive language models suffer exponential decay in correctness over long outputs: if each token has independent error probability e, then (1 − e) n → 0 as n grows. The argument is clean to state and widely cited. It is also brittle, and three lines of recent empirical work make the cracks visible. The first is that only a small subset of tokens—roughly 5% to 10% in the studies that have actually measured it—genuinely depends on long-range context; the rest get more predictable, not less, as context accumulates. The second is geometric: LLM embeddings organize into stratified low-dimensional manifolds, so once a model is working inside one semantic region it tends to stay there even when individual tokens slip. The third concerns what happens when models do err on the consequential tokens—errors turn out to be idiosyncratic across samples rather than systematic, which is why majority-vote ensembles recover so much accuracy. Pulling these together gives a two-rate model, P (correct) ≈ (1−e key ) k ·(1−e non ) n−k , in which k scales sublinearly with n and e non approaches zero with sufficient context. The predicted decay is, at worst, stretched-exponential; often power-law; and when k saturates at some task-specific k max , constant in n. A number of recent capabilities—anchor compression at 99% context reduction, 128K-token retrieval on consumer GPUs, self-consistency gains on reasoning benchmarks—then read as natural consequences of one structural fact rather than independent engineering wins: long-context reliability hinges on a handful of decision points, not on uniform per-token accuracy.

## What we did and found

The paper splits the single per-token error rate of the (1 − e)^n argument in two. Key tokens are the k positions whose correctness depends on long-range context or global knowledge; they fail at a rate e_key. The other n − k fail at a much lower rate, e_non, which falls as context accumulates. Correctness becomes (1 − e_key)^k · (1 − e_non)^(n−k), and the variable that matters is no longer output length but how k grows with n. Each piece of the model is then tied to its own stream of published evidence: measured key-token sparsity, the stratified-manifold geometry of LLM embeddings, and the convergence of correct reasoning paths under sampling. No models are trained and no new experiments are run.

Three regimes follow. Logarithmic growth of k with n gives polynomial decay; fractional-power growth gives stretched-exponential decay; a k that saturates at a task-specific k_max makes reliability constant in n. The evidence leans toward small k. Restricted to key tokens, perplexity tracks downstream performance at Pearson ≈ −0.96, while on the other 91% its correlation is essentially zero. Anchor compression cuts context by 99% with under 1.5% accuracy loss, a result exponential decay has no way to express. Errors correlate, too: minor slips stay on the current manifold, whereas a key-token mistake moves the trajectory onto a wrong one. Because reasoning errors differ from sample to sample, majority voting recovers much of the lost accuracy; on the paper's reading, that is how self-consistency adds 17.9 points on GSM8K.

## Key numbers

| Measure | Value |
|---|---|
| Naive 100-token correctness | 37% |
| Tokens that need long-range context | 9% |
| Key-token perplexity vs. task performance | −0.96 |
| Context reduction under anchor compression | 99% |
| Self-consistency gain on GSM8K | +17.9 points |

## Why it matters

For engineers building long-context or agentic systems, the practical change is the unit of the reliability budget: decisions, not tokens. Paying the quadratic cost of dense attention to recover a sparse signal looks wasteful under the model, so sparse retrieval and anchor compression read as its predictions rather than as lucky engineering. Extra compute belongs at the high-entropy spans where the trajectory can fork, whether that means a tool call fired there, an early exit for tokens the model is sure of, or more sampling exploration where the next token is in doubt. Ensembles pay off on reasoning, where errors are idiosyncratic. On knowledge retrieval they buy little; a missing fact makes the samples fail the same way.

Evaluation shifts with it. Plain perplexity averages over two populations the model handles differently, which is why it predicts task success poorly and why restricting it to key tokens predicts so much better.

The paper is Part 1 of the error trilogy. The Architecture of Errors picks up its central quantity, the number of hard decisions, inside a bounded domain and argues that their failures fall into a small recurring catalogue, so reliability becomes a matter of covering the catalogue rather than outlasting the sequence length. Frontier and Localhost follows the resulting fixes into production systems, where they increasingly live outside the model weights.

## What this does not show

The framework is a synthesis, not a derivation: the paper runs no experiments of its own, and its measured figures come from the cited works. The quantities k, e_key and e_non are observable in principle but have not been measured together on a single benchmark, so the three decay regimes are arguments from supporting evidence rather than fitted curves, and sublinear growth of k is stated as a hypothesis. The geometric account rests on two recent studies (Li and Sarwate; Robinson et al.) that have not been replicated at larger model scales. Nor is there an advance criterion for telling idiosyncratic errors from systematic ones. That distinction is drawn after the fact, from whether ensembling helped.

## Blog post

[Long LLM outputs hinge on a few decisions, not on their length](https://telegrapher.ai/blog/beyond-exponential-decay.md)

## Cite

```bibtex
@misc{arbuzov2026beyond,
  title         = {Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
  year          = {2026},
  note          = {NeurIPS 2026 submission},
  eprint        = {2505.24187},
  archivePrefix = {arXiv},
  url           = {https://telegrapher.ai/research/beyond-exponential-decay/}
}
```
