Long LLM outputs hinge on a few key tokens, so predicted decay is far gentler than (1 − e)^n
Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models
Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets
NeurIPS 2026 submission, May 2026. arXiv:2505.24187
Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length.
What we did and found
The paper splits the single per-token error rate of the (1 − e)^n argument in two. Key tokens are the k positions whose correctness depends on long-range context or global knowledge; they fail at a rate e_key. The other n − k fail at a much lower rate, e_non, which falls as context accumulates. Correctness becomes (1 − e_key)^k · (1 − e_non)^(n−k), and the variable that matters is no longer output length but how k grows with n. Each piece of the model is then tied to its own stream of published evidence: measured key-token sparsity, the stratified-manifold geometry of LLM embeddings, and the convergence of correct reasoning paths under sampling. No models are trained and no new experiments are run.
Three regimes follow. Logarithmic growth of k with n gives polynomial decay; fractional-power growth gives stretched-exponential decay; a k that saturates at a task-specific k_max makes reliability constant in n. The evidence leans toward small k. Restricted to key tokens, perplexity tracks downstream performance at Pearson ≈ −0.96, while on the other 91% its correlation is essentially zero. Anchor compression cuts context by 99% with under 1.5% accuracy loss, a result exponential decay has no way to express. Errors correlate, too: minor slips stay on the current manifold, whereas a key-token mistake moves the trajectory onto a wrong one. Because reasoning errors differ from sample to sample, majority voting recovers much of the lost accuracy; on the paper's reading, that is how self-consistency adds 17.9 points on GSM8K.
Key numbers
| Naive 100-token correctnesswhat (1 − e)^n gives at a 1% per-token error rate; the prediction the paper argues against | 37% |
| Tokens that need long-range contextshare of tokens in natural text scoring LSD > 2 (Fang et al., 2024); adversarial-perturbation work lands in the same 5%–10% band | 9% |
| Key-token perplexity vs. task performancePearson correlation when perplexity is restricted to key tokens (LongPPL); on the other 91% of tokens it is essentially zero | −0.96 |
| Context reduction under anchor compressionwith < 1.5% accuracy loss (Pang et al., 2024), read as evidence that k is bounded for those tasks | 99% |
| Self-consistency gain on GSM8Kmajority vote over sampled reasoning paths, no retraining (Wang et al., 2023) | +17.9 points |
What this does not show
The framework is a synthesis, not a derivation: the paper runs no experiments of its own, and its measured figures come from the cited works. The quantities k, e_key and e_non are observable in principle but have not been measured together on a single benchmark, so the three decay regimes are arguments from supporting evidence rather than fitted curves, and sublinear growth of k is stated as a hypothesis. The geometric account rests on two recent studies (Li and Sarwate; Robinson et al.) that have not been replicated at larger model scales. Nor is there an advance criterion for telling idiosyncratic errors from systematic ones. That distinction is drawn after the fact, from whether ensembling helped.
The authors' abstract
A common pessimistic argument holds that autoregressive language models suffer exponential decay in correctness over long outputs: if each token has independent error probability e, then (1 − e) n → 0 as n grows. The argument is clean to state and widely cited. It is also brittle, and three lines of recent empirical work make the cracks visible. The first is that only a small subset of tokens—roughly 5% to 10% in the studies that have actually measured it—genuinely depends on long-range context; the rest get more predictable, not less, as context accumulates. The second is geometric: LLM embeddings organize into stratified low-dimensional manifolds, so once a model is working inside one semantic region it tends to stay there even when individual tokens slip. The third concerns what happens when models do err on the consequential tokens—errors turn out to be idiosyncratic across samples rather than systematic, which is why majority-vote ensembles recover so much accuracy. Pulling these together gives a two-rate model, P (correct) ≈ (1−e key ) k ·(1−e non ) n−k , in which k scales sublinearly with n and e non approaches zero with sufficient context. The predicted decay is, at worst, stretched-exponential; often power-law; and when k saturates at some task-specific k max , constant in n. A number of recent capabilities—anchor compression at 99% context reduction, 128K-token retrieval on consumer GPUs, self-consistency gains on reasoning benchmarks—then read as natural consequences of one structural fact rather than independent engineering wins: long-context reliability hinges on a handful of decision points, not on uniform per-token accuracy.
The rest of the trilogy
Cite
@misc{arbuzov2026beyond,
title = {Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models},
author = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
year = {2026},
note = {NeurIPS 2026 submission},
eprint = {2505.24187},
archivePrefix = {arXiv},
url = {https://telegrapher.ai/research/beyond-exponential-decay/}
}Builds on
- Fang et al. (2024). What is wrong with perplexity for long-context language modeling?
- Li and Sarwate (2025). Unraveling the localized latents: Learning stratified manifold structures in llm embedding space with sparse mixture-of-experts.
- Wang et al. (2023). Self-consistency improves chain of thought reasoning in language models.
- Pang et al. (2024). Anchor-based large language models.
- Liu et al. (2024). Retrievalattention: Accelerating long-context llm inference via vector retrieval.