---
type: post
title: Long LLM outputs hinge on a few decisions, not on their length
date: '2026-10-09'
description: About 9% of tokens carry long-range dependency and errors cluster, so
  the predicted decay of long LLM outputs is far gentler than (1 − e)^n.
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Ziwei Dong
- Dmitri Kalaev
- Alexey Shvets
paper: https://telegrapher.ai/research/beyond-exponential-decay.md
html: https://telegrapher.ai/blog/beyond-exponential-decay/
---

# Long LLM outputs hinge on a few decisions, not on their length

Here is the arithmetic behind a familiar worry about autoregressive language models. Give each token a 1% chance of going wrong, independently of the others. A 100-token answer is then correct (0.99)^100 of the time, which comes to about 37%. Make the answer longer and the figure slides toward zero. LeCun applied the argument to LLMs in 2023, and it is widely cited as a reason such models must hit a wall.

The arithmetic is sound when its assumptions hold. What it predicts is not what we see. Outputs that run for pages hold together, and models routinely revise an earlier reading halfway through. In [the paper](/research/beyond-exponential-decay/) we argue that the trouble is not the algebra but the population it averages over.

## Most tokens are not where long outputs break

Take a story that has already mentioned a bridge closure and a ferry and now reaches this sentence: "Since the bridge closed on Monday, she took the ferry instead." By now most of its words are close to forced. A few tokens are different. "Monday" is a factual commitment that has to agree with something said earlier. "Since" and "instead" are logical operators: they assert a cause and a substitution. "She" has to point at the right person. Get any of those wrong and the sentence is still fluent. It is also false.

Tokens like these are key tokens: factual claims, logical operators, points of co-reference, transitions between topics. Whether they come out right depends on long-range context or on knowledge from outside the text. The rest, the non-key tokens, are held in place by local syntax, common word pairings and whatever the text has established. A sentence picked to display key tokens is crowded with them; in ordinary text they are a small minority.

So we split the single error rate in two. With *k* key tokens in an output of *n*:

`P(correct) ≈ (1 − e_key)^k · (1 − e_non)^(n − k)`

Here `e_key` is the error rate on key tokens and `e_non` the much smaller rate on everything else. Because `e_non` keeps falling as context accumulates, later non-key tokens carry less risk than earlier ones, and output length stops being the main variable. What decides reliability is how *k* grows with *n*.

## Slips stay local; wrong commitments travel

The (1 − e)^n form also assumes errors at different positions are independent. They are not, and on our account the reason is geometric. Li and Sarwate (2025) probed LLM embedding spaces and found a stratified manifold: a union of low-dimensional regions, each aligned with a semantic domain. Picture the hidden state moving inside one region, set by the context so far. A synonym swap or a grammatical wobble stays inside the region and leaves nothing downstream. A key-token error does something else. It moves the trajectory into a different region, and what follows stays fluent and consistent with the wrong commitment.

Robinson et al. (2025) approach the geometry from individual tokens: perturbing one inside a region barely changes generation, while perturbing one at a junction between regions reroutes it. Junctions are where key tokens should sit.

Mistakes bunch at a few junctions instead of spreading out. With disruptive errors confined to the *k* key tokens, the chance of any one occurring is at most *k* times `e_key`, well below the 1 − (1 − e)^n failure rate of the original argument when *k* is much smaller than *n*.

Ensembles fit the same picture. When key-token errors differ from sample to sample, paths go wrong at different junctions and a majority vote can still land on the right answer. Self-consistency (Wang et al., 2023) gained 17.9 points on GSM8K with no retraining. On our reading, the gain comes from correct paths converging while wrong ones scatter. But when the error is systematic, as with a fact the model simply lacks, the samples fail in the same way and voting has nothing to recover.

## The curve depends on how decisions grow with length

Try different growth rates for *k* in the two-rate model and the curve changes shape:

| How errors are modeled | Predicted reliability |
|---|---|
| One independent rate for all tokens (the original argument) | exponential decay |
| *k* grows like log *n* | polynomial (power-law) decay |
| *k* grows as a fractional power of *n* | stretched-exponential decay |
| *k* saturates at a task-specific `k_max` | constant in *n* |

We hypothesize that *k* grows sublinearly and may saturate, since even a book-length argument turns on a bounded number of critical decisions. The published measurements bear most directly on a simpler point: *k* is small next to *n*. Fang et al. (2024) put the share of tokens in natural text that depend meaningfully on distant context at about 9%. Adversarial-perturbation studies, built for a different purpose, land close to that figure: flip a small set of well-chosen tokens and the model's decision changes, while randomly perturbing far more does nothing. Restrict perplexity to Fang et al.'s key tokens and it tracks downstream task performance closely; on the remaining tokens the correlation is near zero.

RetrievalAttention (Liu et al., 2024) shows the same sparsity from the systems side. Llama-3 puts nearly all its attention mass on roughly a thousand tokens out of a hundred thousand, so the method restricts attention to those tokens and fits long contexts on a single consumer GPU. That is *k* ≪ *n* put to work.

The saturating regime is the strange one, and it has a witness. Anchor compression (Pang et al., 2024) cuts context by 99% with under 1.5% accuracy loss. An exponential curve has no room for that result. Under the two-rate model it means *k* is effectively bounded for those tasks.

## Spend compute where the trajectory can fork

If reliability is decided at a few junctions, the budget should go there too. Paying the quadratic cost of dense attention to recover a sparse signal looks wasteful, and the sparse methods that skip it look predictable rather than lucky. Extra compute belongs at high-entropy spans — where the next token is in doubt. Tool-integrated reasoning already fires code execution there rather than uniformly, and early-exit and adaptive-temperature methods make the same split token by token. Ensembles belong on reasoning, where errors vary from sample to sample, not on knowledge retrieval, where they repeat.

Evaluation changes with it. Perplexity averaged over the whole output mixes two populations the model handles differently, which is why the key-token version predicts better. In the success-plateau curves of Costello et al. (2025), reliability does not slide smoothly either: long flat stretches end in sharp drops, a staircase rather than an exponential.

## Parts 2 and 3: a catalogue of failures, and fixes outside the weights

This paper opens a three-part argument about errors. [The Architecture of Errors](/research/architecture-of-errors/) takes the hard decisions inside a bounded domain, such as legal review or code repair, and argues that their failures fall into a small recurring catalogue, so reliability becomes a matter of covering that catalogue rather than outlasting the sequence length. [Frontier and Localhost](/research/frontier-and-localhost/) follows the fixes into production systems, where they increasingly live outside the model weights, and asks what a disciplined optimizer for that layer would look like.

## What we have not measured

This is a synthesis rather than a derivation. We ran no experiments of our own, and the measured figures above come from the work we cite. The three quantities the model needs, *k*, `e_key` and `e_non`, are observable in principle but have not been measured together on a single benchmark. The regimes in the table are therefore argued from evidence, not fitted to data, and sublinear growth of *k* remains a hypothesis. The geometric account leans on two recent studies, Li and Sarwate (2025) and Robinson et al. (2025); replication at larger model scales would tighten it. Nor do we have an advance rule for telling idiosyncratic errors from systematic ones; for now we learn a task's regime after the fact, from whether ensembling helped.

The obvious next step is to measure *k*(*n*) on a fixed benchmark. Two more stay open: bounds drawn from attention statistics instead of union-bound heuristics, and a bridge from token-level decay to the step-level decay that matters for agents. If *k* saturates on the tasks people care about, then length was the wrong thing to count. The number of decisions was.
