{
  "slug": "beyond-exponential-decay",
  "title": "Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models",
  "short": "BED",
  "line": "errors",
  "line_name": "Error-accumulation",
  "part": 1,
  "status": "NeurIPS 2026 submission",
  "date": "2026-05-04",
  "authors": [
    "Mikhail L Arbuzov",
    "Sisong Bei",
    "Ziwei Dong",
    "Dmitri Kalaev",
    "Alexey Shvets"
  ],
  "abstract": "A common pessimistic argument holds that autoregressive language models suffer exponential decay in correctness over long outputs: if each token has independent error probability e, then (1 − e) n → 0 as n grows. The argument is clean to state and widely cited. It is also brittle, and three lines of recent empirical work make the cracks visible. The first is that only a small subset of tokens—roughly 5% to 10% in the studies that have actually measured it—genuinely depends on long-range context; the rest get more predictable, not less, as context accumulates. The second is geometric: LLM embeddings organize into stratified low-dimensional manifolds, so once a model is working inside one semantic region it tends to stay there even when individual tokens slip. The third concerns what happens when models do err on the consequential tokens—errors turn out to be idiosyncratic across samples rather than systematic, which is why majority-vote ensembles recover so much accuracy. Pulling these together gives a two-rate model, P (correct) ≈ (1−e key ) k ·(1−e non ) n−k , in which k scales sublinearly with n and e non approaches zero with sufficient context. The predicted decay is, at worst, stretched-exponential; often power-law; and when k saturates at some task-specific k max , constant in n. A number of recent capabilities—anchor compression at 99% context reduction, 128K-token retrieval on consumer GPUs, self-consistency gains on reasoning benchmarks—then read as natural consequences of one structural fact rather than independent engineering wins: long-context reliability hinges on a handful of decision points, not on uniform per-token accuracy.",
  "tldr": "Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length.",
  "pages": 15,
  "html": "https://telegrapher.ai/research/beyond-exponential-decay/",
  "md": "https://telegrapher.ai/research/beyond-exponential-decay.md",
  "reader": "https://telegrapher.ai/research/beyond-exponential-decay/read/",
  "pdf": "https://telegrapher.ai/papers/beyond-exponential-decay/beyond-exponential-decay.pdf",
  "arxiv": "2505.24187",
  "openreview": "https://openreview.net/forum?id=cZl2t7D4m6",
  "post": "https://telegrapher.ai/blog/beyond-exponential-decay/",
  "bibtex": "@misc{arbuzov2026beyond,\n  title         = {Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models},\n  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},\n  year          = {2026},\n  note          = {NeurIPS 2026 submission},\n  eprint        = {2505.24187},\n  archivePrefix = {arXiv},\n  url           = {https://telegrapher.ai/research/beyond-exponential-decay/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "Long LLM outputs hinge on a few key tokens, so predicted decay is far gentler than (1 − e)^n"
    },
    {
      "label": "TL;DR",
      "text": "Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length."
    },
    {
      "label": "Method",
      "text": "A synthesis with no new experiments: the paper folds published measurements of key-token sparsity, embedding geometry and ensemble convergence into a two-rate error model, then re-reads existing systems results (anchor compression, RetrievalAttention, self-consistency, tool-integrated reasoning) through it."
    },
    {
      "label": "Key result",
      "text": "Naive 100-token correctness: 37%; Tokens that need long-range context: 9%; Key-token perplexity vs. task performance: −0.96"
    },
    {
      "label": "Why it matters",
      "text": "For engineers building long-context or agentic systems, the practical change is the unit of the reliability budget: decisions, not tokens."
    },
    {
      "label": "Limits",
      "text": "The framework is a synthesis, not a derivation: the paper runs no experiments of its own, and its measured figures come from the cited works."
    },
    {
      "label": "Status",
      "text": "NeurIPS 2026 submission, May 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/beyond-exponential-decay/read/, PDF /papers/beyond-exponential-decay/beyond-exponential-decay.pdf, arXiv:2505.24187 https://arxiv.org/abs/2505.24187"
    }
  ],
  "note": {
    "slug": "beyond-exponential-decay",
    "claim_title": "Long LLM outputs hinge on a few key tokens, so predicted decay is far gentler than (1 − e)^n",
    "meta_description": "Why (1 − e)^n misdescribes long LLM outputs: about 9% of tokens carry long-range dependency, errors correlate, and predicted decay is far gentler.",
    "tldr": "Treat every token as an equal, independent chance to fail and long outputs look doomed. Published measurements find about 9% of tokens depend on long-range context, and errors correlate, so reliability tracks key decisions, not output length.",
    "gist": "A widely cited argument says autoregressive models must fail on long outputs. With per-token error e, correctness falls as (1 − e)^n; at a 1% error rate a 100-token chain is right about 37% of the time. The paper's reply is that the formula averages over two different populations. Key tokens (factual claims, logical operators, points of co-reference, topic transitions) depend on long-range context, and published measurements put their share near 9%, inside a 5%–10% band that recurs across methods. The rest are pinned down by local syntax and accumulated context, so their error rate falls toward zero as the text grows. Errors also correlate. A minor slip stays inside one semantic region of the embedding space; a wrong commitment carries the continuation into another, where it stays fluent and wrong. Depending on how the number of key tokens grows with length, the resulting two-rate model predicts power-law, stretched-exponential or constant reliability. It is a synthesis of published results, and its parameters have not yet been measured together on a single benchmark.",
    "method": "A synthesis with no new experiments: the paper folds published measurements of key-token sparsity, embedding geometry and ensemble convergence into a two-rate error model, then re-reads existing systems results (anchor compression, RetrievalAttention, self-consistency, tool-integrated reasoning) through it.",
    "summary_html": [
      "The paper splits the single per-token error rate of the <code>(1 − e)^n</code> argument in two. Key tokens are the <em>k</em> positions whose correctness depends on long-range context or global knowledge; they fail at a rate <code>e_key</code>. The other <em>n − k</em> fail at a much lower rate, <code>e_non</code>, which falls as context accumulates. Correctness becomes <code>(1 − e_key)^k · (1 − e_non)^(n−k)</code>, and the variable that matters is no longer output length but how <em>k</em> grows with <em>n</em>. Each piece of the model is then tied to its own stream of published evidence: measured key-token sparsity, the stratified-manifold geometry of LLM embeddings, and the convergence of correct reasoning paths under sampling. No models are trained and no new experiments are run.",
      "Three regimes follow. Logarithmic growth of <em>k</em> with <em>n</em> gives polynomial decay; fractional-power growth gives stretched-exponential decay; a <em>k</em> that saturates at a task-specific <code>k_max</code> makes reliability constant in <em>n</em>. The evidence leans toward small <em>k</em>. Restricted to key tokens, perplexity tracks downstream performance at Pearson ≈ −0.96, while on the other 91% its correlation is essentially zero. Anchor compression cuts context by 99% with under 1.5% accuracy loss, a result exponential decay has no way to express. Errors correlate, too: minor slips stay on the current manifold, whereas a key-token mistake moves the trajectory onto a wrong one. Because reasoning errors differ from sample to sample, majority voting recovers much of the lost accuracy; on the paper's reading, that is how self-consistency adds 17.9 points on GSM8K."
    ],
    "key_numbers": [
      {
        "label": "Naive 100-token correctness",
        "value": "37%",
        "context": "what (1 − e)^n gives at a 1% per-token error rate; the prediction the paper argues against",
        "bad": true
      },
      {
        "label": "Tokens that need long-range context",
        "value": "9%",
        "context": "share of tokens in natural text scoring LSD > 2 (Fang et al., 2024); adversarial-perturbation work lands in the same 5%–10% band"
      },
      {
        "label": "Key-token perplexity vs. task performance",
        "value": "−0.96",
        "context": "Pearson correlation when perplexity is restricted to key tokens (LongPPL); on the other 91% of tokens it is essentially zero"
      },
      {
        "label": "Context reduction under anchor compression",
        "value": "99%",
        "context": "with < 1.5% accuracy loss (Pang et al., 2024), read as evidence that k is bounded for those tasks"
      },
      {
        "label": "Self-consistency gain on GSM8K",
        "value": "+17.9 points",
        "context": "majority vote over sampled reasoning paths, no retraining (Wang et al., 2023)"
      }
    ],
    "editorial_html": [
      "For engineers building long-context or agentic systems, the practical change is the unit of the reliability budget: decisions, not tokens. Paying the quadratic cost of dense attention to recover a sparse signal looks wasteful under the model, so sparse retrieval and anchor compression read as its predictions rather than as lucky engineering. Extra compute belongs at the high-entropy spans where the trajectory can fork, whether that means a tool call fired there, an early exit for tokens the model is sure of, or more sampling exploration where the next token is in doubt. Ensembles pay off on reasoning, where errors are idiosyncratic. On knowledge retrieval they buy little; a missing fact makes the samples fail the same way.",
      "Evaluation shifts with it. Plain perplexity averages over two populations the model handles differently, which is why it predicts task success poorly and why restricting it to key tokens predicts so much better.",
      "The paper is Part 1 of the error trilogy. <em>The Architecture of Errors</em> picks up its central quantity, the number of hard decisions, inside a bounded domain and argues that their failures fall into a small recurring catalogue, so reliability becomes a matter of covering the catalogue rather than outlasting the sequence length. <em>Frontier and Localhost</em> follows the resulting fixes into production systems, where they increasingly live outside the model weights."
    ],
    "limitations": "The framework is a synthesis, not a derivation: the paper runs no experiments of its own, and its measured figures come from the cited works. The quantities k, e_key and e_non are observable in principle but have not been measured together on a single benchmark, so the three decay regimes are arguments from supporting evidence rather than fitted curves, and sublinear growth of k is stated as a hypothesis. The geometric account rests on two recent studies (Li and Sarwate; Robinson et al.) that have not been replicated at larger model scales. Nor is there an advance criterion for telling idiosyncratic errors from systematic ones. That distinction is drawn after the fact, from whether ensembling helped.",
    "concept_terms": [
      "two-rate error model",
      "key tokens",
      "exponential error accumulation",
      "stratified manifold",
      "self-consistency",
      "LongPPL"
    ],
    "references": [
      "Fang et al. (2024). What is wrong with perplexity for long-context language modeling?",
      "Li and Sarwate (2025). Unraveling the localized latents: Learning stratified manifold structures in llm embedding space with sparse mixture-of-experts.",
      "Wang et al. (2023). Self-consistency improves chain of thought reasoning in language models.",
      "Pang et al. (2024). Anchor-based large language models.",
      "Liu et al. (2024). Retrievalattention: Accelerating long-context llm inference via vector retrieval."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/beyond-exponential-decay.txt",
    "verification": {
      "claims": [
        {
          "claim": "The (1 − e)^n argument that long outputs must fail is widely cited (gist)",
          "verdict": "CONFIRMED",
          "source_quote": "if each token has independent error probability e, then (1 − e) n → 0 as n grows. The argument is clean to state and widely cited."
        },
        {
          "claim": "At a 1% per-token error rate a 100-token chain is correct about 37% of the time; key number 37% is the naive prediction the paper argues against",
          "verdict": "CONFIRMED",
          "source_quote": "If each token has even a 1% error rate then a 100-token chain is correct only (0.99) 100 ≈ 37% of the time"
        },
        {
          "claim": "The formula averages over two different populations of tokens",
          "verdict": "CONFIRMED",
          "source_quote": "the independent-error model fails because it averages over a heterogeneous population of tokens"
        },
        {
          "claim": "Key tokens are factual claims, logical operators, points of co-reference and topic transitions, whose correctness depends on long-range context or global knowledge",
          "verdict": "CONFIRMED",
          "source_quote": "Key tokens are the ones whose correctness genuinely depends on long-range context or global knowledge—factual claims, logical operators, points of co-reference, transitions between topics."
        },
        {
          "claim": "About 9% of tokens in natural text depend on long-range context, scoring LSD > 2 (Fang et al., 2024); key number 9%",
          "verdict": "CONFIRMED",
          "source_quote": "whose long-short difference (LSD) metric finds only ∼ 9% of tokens in natural text scoring LSD> 2"
        },
        {
          "claim": "The share sits in a 5%–10% band that recurs across methods; adversarial-perturbation work lands in the same band",
          "verdict": "CONFIRMED",
          "source_quote": "Morris et al. [2022] flip model decisions by perturbing 5%–10% of strategically chosen tokens ... The 5%–10% figure recurs across methodologies built to measure different things"
        },
        {
          "claim": "Non-key tokens are pinned down by local syntax and accumulated context, and their error rate falls toward zero",
          "verdict": "CONFIRMED",
          "source_quote": "The rest are constrained by local syntax and the accumulating context, and their error rate goes to zero rather than to a constant."
        },
        {
          "claim": "Errors are not independent across positions; they correlate",
          "verdict": "CONFIRMED",
          "source_quote": "The (1 − e) n form also assumes errors are independent across positions. They are not."
        },
        {
          "claim": "Minor slips stay on the current manifold; a key-token mistake moves the trajectory onto a different one, and the continuation is fluent but wrong",
          "verdict": "CONFIRMED",
          "source_quote": "a momentary wobble that stays on M C and has no downstream effect ... after which subsequent tokens cohere with the wrong commitment. The result is a fluent-but-wrong continuation."
        },
        {
          "claim": "LLM embeddings have a stratified-manifold geometry (Li and Sarwate, 2025)",
          "verdict": "CONFIRMED",
          "source_quote": "a union of low-dimensional submanifolds aligned with semantic domain"
        },
        {
          "claim": "Two-rate model: P(correct) ≈ (1 − e_key)^k · (1 − e_non)^(n−k), with e_non much lower than e_key",
          "verdict": "CONFIRMED",
          "source_quote": "P (correct) ≈ (1 − e key ) k · (1 − e non ) n−k ... e non the much lower error rate for the remaining n − k"
        },
        {
          "claim": "Logarithmic growth of k gives polynomial (power-law) decay, fractional-power growth gives stretched-exponential decay, and saturation at k_max makes reliability constant in n",
          "verdict": "CONFIRMED",
          "source_quote": "Logarithmic key-token growth (k ∼ log n): polynomial decay n −c ... stretched-exponential decay ... Saturating (k → k max ): reliability becomes constant in n ... The predicted decay is, at worst, stretched-exponential; often power-law"
        },
        {
          "claim": "The predicted decay is far gentler than (1 − e)^n (claim_title, meta_description; corrected from an unhedged 'decays' / 'decay turns power-law or flat')",
          "verdict": "CONFIRMED",
          "source_quote": "the predicted decay is far gentler than (1 − e) n"
        },
        {
          "claim": "Reliability tracks a few key decisions, not output length; the variable that matters is how k grows with n",
          "verdict": "CONFIRMED",
          "source_quote": "long-context reliability hinges on a handful of decision points, not on uniform per-token accuracy ... Reliability depends on k key decisions, not on n tokens"
        },
        {
          "claim": "The paper is a synthesis: no new experiments, no models trained; its measured figures come from the cited works",
          "verdict": "CONFIRMED",
          "source_quote": "it runs no new experiments and trains no new models ... The paper makes no empirical claims of its own"
        },
        {
          "claim": "Each piece of the model is tied to published evidence on key-token sparsity, embedding geometry and ensemble convergence",
          "verdict": "CONFIRMED",
          "source_quote": "attention sparsity, embedding geometry, and ensemble convergence—that supports each component of the model"
        },
        {
          "claim": "The paper re-reads anchor compression, RetrievalAttention, self-consistency and tool-integrated reasoning through the model",
          "verdict": "CONFIRMED",
          "source_quote": "(anchor compression, retrieval-augmented attention, self-consistency, tool integration) ... RetrievalAttention exploits this directly"
        },
        {
          "claim": "The evidence leans toward small k",
          "verdict": "CONFIRMED",
          "source_quote": "The framework says k/n is small and stable"
        },
        {
          "claim": "Perplexity restricted to key tokens (LongPPL) tracks downstream performance at Pearson ≈ −0.96; key number −0.96",
          "verdict": "CONFIRMED",
          "source_quote": "perplexity restricted to those tokens correlates with downstream task performance at Pearson ≈ −0.96 ... LongPPL restricts perplexity to key tokens"
        },
        {
          "claim": "On the other 91% of tokens the correlation with task success is essentially zero",
          "verdict": "CONFIRMED",
          "source_quote": "Perplexity on the other 91% tracks task success at essentially zero correlation."
        },
        {
          "claim": "Anchor compression cuts context by 99% with under 1.5% accuracy loss (Pang et al., 2024); key number 99%",
          "verdict": "CONFIRMED",
          "source_quote": "Pang et al. [2024] achieved 99% context reduction with < 1.5% accuracy loss"
        },
        {
          "claim": "Exponential decay cannot express the anchor-compression result, which reads as evidence that k is bounded for those tasks",
          "verdict": "CONFIRMED",
          "source_quote": "a result that requires k to be effectively bounded for the relevant tasks. There is no way to express that under exponential decay"
        },
        {
          "claim": "Self-consistency (majority vote over sampled reasoning paths) adds +17.9 points on GSM8K with no retraining (Wang et al., 2023); key number +17.9 points",
          "verdict": "CONFIRMED",
          "source_quote": "sample multiple reasoning paths, take the majority answer. The gain was +17.9 points on GSM8K, with no retraining."
        },
        {
          "claim": "On the paper's reading, the self-consistency gain comes from reasoning errors differing across samples, so majority voting recovers much of the lost accuracy",
          "verdict": "CONFIRMED",
          "source_quote": "sit closer to the ρ = 0 regime than to ρ = 1, which is why a method that adds no parameters and no training data can lift GSM8K accuracy by 17.9 points ... majority-vote ensembles recover so much accuracy"
        },
        {
          "claim": "Ensembles pay off on reasoning, not on knowledge retrieval, where a missing fact makes the samples fail the same way",
          "verdict": "CONFIRMED",
          "source_quote": "knowledge gaps produce systematic errors (every sample fails the same way); reasoning slips produce idiosyncratic ones (samples fail differently)"
        },
        {
          "claim": "Dense attention pays a quadratic cost to recover a sparse signal, so sparse retrieval and anchor compression read as predictions of the model",
          "verdict": "CONFIRMED",
          "source_quote": "These methods stop being clever tricks and become predictable: dense attention pays a quadratic cost to recover a sparse signal."
        },
        {
          "claim": "Extra compute belongs at high-entropy spans: tool calls there, early exit for confident tokens, more sampling exploration at uncertain tokens",
          "verdict": "CONFIRMED",
          "source_quote": "tool-integrated reasoning fires Python execution at high-entropy spans rather than uniformly ... lets confident tokens skip layers altogether ... raises exploration at uncertain tokens and damps it elsewhere"
        },
        {
          "claim": "Plain perplexity mixes two populations, which is why it predicts task success poorly and the key-token version predicts better",
          "verdict": "CONFIRMED",
          "source_quote": "Plain perplexity averages over a population the model is trying to handle separately, which is why it underperforms as a predictor."
        },
        {
          "claim": "k, e_key and e_non have not been measured together on a single benchmark; the decay regimes are arguments from supporting evidence, not fitted curves",
          "verdict": "CONFIRMED",
          "source_quote": "observable in principle but not yet jointly measured on a single benchmark, so the predicted decay regimes are arguments from supporting evidence rather than fitted curves"
        },
        {
          "claim": "Sublinear growth of k is stated as a hypothesis",
          "verdict": "CONFIRMED",
          "source_quote": "We hypothesize that k grows sublinearly with n"
        },
        {
          "claim": "The geometric account rests on Li and Sarwate and Robinson et al., both recent and not replicated at larger model scales",
          "verdict": "CONFIRMED",
          "source_quote": "both recent and not yet replicated at larger scales"
        },
        {
          "claim": "There is no advance criterion for telling idiosyncratic from systematic errors; the distinction is drawn after the fact from whether ensembling helped",
          "verdict": "CONFIRMED",
          "source_quote": "no quantitative criterion for distinguishing the idiosyncratic from the systematic regime in advance, only the post-hoc observation that ensemble methods help in one and not the other"
        },
        {
          "claim": "Reference: Fang et al. (2024) is in the bibliography",
          "verdict": "CONFIRMED",
          "source_quote": "What is wrong with perplexity for long-context language modeling?"
        },
        {
          "claim": "Reference: Li and Sarwate (2025) is in the bibliography",
          "verdict": "CONFIRMED",
          "source_quote": "Unraveling the localized latents: Learning stratified manifold structures in llm embedding space with sparse mixture-of-experts."
        },
        {
          "claim": "Reference: Wang et al. (2023) is in the bibliography",
          "verdict": "CONFIRMED",
          "source_quote": "Self-consistency improves chain of thought reasoning in language models."
        },
        {
          "claim": "Reference: Pang et al. (2024) is in the bibliography",
          "verdict": "CONFIRMED",
          "source_quote": "Anchor-based large language models."
        },
        {
          "claim": "Reference: Liu et al. (2024) is in the bibliography",
          "verdict": "CONFIRMED",
          "source_quote": "Retrievalattention: Accelerating long-context llm inference via vector retrieval."
        },
        {
          "claim": "Editorial: the paper is Part 1 of the error trilogy (site metadata; the paper text itself does not mention the trilogy)",
          "verdict": "CONFIRMED",
          "source_quote": "[content/papers.yaml] slug: beyond-exponential-decay short: BED ... line: errors part: 1"
        },
        {
          "claim": "Editorial: The Architecture of Errors takes up the number of hard decisions inside bounded domains and argues failures fall into a small recurring catalogue, making reliability a coverage problem rather than a sequence-length one (checked against that paper's abstract, not this paper's text)",
          "verdict": "CONFIRMED",
          "source_quote": "[content/abstracts.yaml] where the number of hard decisions itself grows with task length ... failures are sparse, repetitive, and concentrated in a small recurring catalogue, so reliability becomes a local catalogue-discovery and intervention-coverage problem rather than an exponential token-length problem"
        },
        {
          "claim": "Editorial: Frontier and Localhost follows fixes into production systems that increasingly adapt outside the model weights (checked against that paper's abstract, not this paper's text)",
          "verdict": "CONFIRMED",
          "source_quote": "[content/abstracts.yaml] Production LLM systems increasingly adapt outside the model weights."
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}