{
  "slug": "architecture-of-errors",
  "title": "The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability",
  "short": "AoE",
  "line": "errors",
  "line_name": "Error-accumulation",
  "part": 2,
  "status": "COLM 2026 workshop poster",
  "date": "2026-08-04",
  "authors": [
    "Mikhail L Arbuzov",
    "Sisong Bei",
    "Ziwei Dong",
    "Dmitri Kalaev",
    "Alexey Shvets"
  ],
  "abstract": "Reliability is the implicit goal of context management — what a model retrieves, remembers, and is scaffolded with is chosen so it fails less — yet it is usually analysed asymptotically, as if each token compounded the risk. We argue the object to track is not raw sequence length but a small, local catalogue of recurring failure modes. Universal LLM reliability is not a finite-library problem: across all possible tasks, tools, schemas, knowledge sources, and evaluator expectations, new intervention-distinguishable failure modes can appear without bound, so no finite intervention dictionary can guarantee bounded residual error for every such mode. But deployed systems do not operate over the whole universe. They operate inside operationally bounded patches (legal review, medical RAG, code repair, customer-support agents, contract extraction) with recurring tasks, schemas, tools, and evaluator expectations—the operational envelope that the context scaffold of retrieval, memory, tools, and orchestration defines and maintains. Within such patches, empirical evidence suggests failures are sparse, repetitive, and concentrated in a small recurring catalogue, so reliability becomes a local catalogue-discovery and intervention-coverage problem rather than an exponential token-length problem. We formalize this transition with two propositions and one corollary. Proposition 1 is the worst-case-mode-wise negative result: no finite intervention dictionary covers every distinguishable failure mode of an unbounded domain. Corollary 1 is the inverse-discovery implication: the logarithmic upper bound on mode discovery cannot accommodate linearly more distinct tail modes without exponentially more observed hard-failure events. Proposition 2 is the positive patch-local result: under log active-mode exposure and headheavy coverage, a sufficient per-hard-decision intervention budget grows polylogarithmically in sequence length and becomes domain-constant once the patch catalogue saturates. The framework relocates rather than dissolves long-context difficulty: where the number of hard decisions itself grows with task length, reliability remains hard; the contribution is to identify the on-axis intervention rather than to make those regimes easy.",
  "tldr": "No finite fix list covers every failure mode of open-ended LLM use. Inside one deployment, published taxonomies suggest failures recur in a small catalogue, so a sufficient fix library grows slowly with sequence length, then levels off.",
  "pages": 24,
  "html": "https://telegrapher.ai/research/architecture-of-errors/",
  "md": "https://telegrapher.ai/research/architecture-of-errors.md",
  "reader": "https://telegrapher.ai/research/architecture-of-errors/read/",
  "pdf": "https://telegrapher.ai/papers/architecture-of-errors/architecture-of-errors.pdf",
  "arxiv": "2605.30628",
  "openreview": "https://openreview.net/forum?id=DimJCP8Bz1",
  "post": "https://telegrapher.ai/blog/architecture-of-errors/",
  "bibtex": "@misc{arbuzov2026the,\n  title         = {The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability},\n  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},\n  year          = {2026},\n  note          = {COLM 2026 workshop poster},\n  eprint        = {2605.30628},\n  archivePrefix = {arXiv},\n  url           = {https://telegrapher.ai/research/architecture-of-errors/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "No finite fix list covers open-ended LLM use; inside a deployment, a sufficient list levels off"
    },
    {
      "label": "TL;DR",
      "text": "No finite fix list covers every failure mode of open-ended LLM use. Inside one deployment, published taxonomies suggest failures recur in a small catalogue, so a sufficient fix library grows slowly with sequence length, then levels off."
    },
    {
      "label": "Method",
      "text": "Theory plus synthesis: two propositions and a corollary about intervention libraries, with the mode-discovery rate calibrated on endpoint counts from three published failure taxonomies (ErrorAtlas, a HumanEval error categorisation, MWPES-300K), and the framework checked against about 60 published results, including 28 intervention citations across six capability axes and re-audits of five steep-decay papers."
    },
    {
      "label": "Key result",
      "text": "HumanEval failures from two error types: 86.35%; Failure categories in ErrorAtlas: 17; Extra failures needed to find five more modes: ≈15×"
    },
    {
      "label": "Why it matters",
      "text": "For a team running a model inside one job, contract extraction say, or code repair, the unit of reliability work moves from the model to the patch."
    },
    {
      "label": "Limits",
      "text": "The logarithmic discovery rate is a postulate, not a theorem."
    },
    {
      "label": "Status",
      "text": "COLM 2026 workshop poster, August 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/architecture-of-errors/read/, PDF /papers/architecture-of-errors/architecture-of-errors.pdf, arXiv:2605.30628 https://arxiv.org/abs/2605.30628"
    }
  ],
  "note": {
    "slug": "architecture-of-errors",
    "claim_title": "No finite fix list covers open-ended LLM use; inside a deployment, a sufficient list levels off",
    "meta_description": "Why no finite list of fixes can cover open-ended LLM use, and why recurring failures inside one deployment keep a sufficient intervention library small.",
    "tldr": "No finite fix list covers every failure mode of open-ended LLM use. Inside one deployment, published taxonomies suggest failures recur in a small catalogue, so a sufficient fix library grows slowly with sequence length, then levels off.",
    "gist": "The usual reliability worry about language models is length: if every token can fail, long outputs should collapse. This paper tracks a different object, the catalogue of recurring failure modes. Across open-ended use that catalogue can grow without bound. Proposition 1 shows that if a domain keeps producing failures that escape any finite set of fixes for the earlier ones, no finite intervention dictionary can hold every failure mode below a fixed error tolerance. Deployed systems, though, run inside bounded patches such as legal review or code repair, where tasks, schemas and evaluators recur, and the published taxonomies find small, head-heavy catalogues. If the modes a sequence can activate grow logarithmically and coverage is head-heavy, Proposition 2 gives a sufficient per-hard-decision intervention budget that grows polylogarithmically with sequence length and stops growing once the patch catalogue saturates. Corollary 1 prices the tail: the minimum number of observed failures grows exponentially with the number of modes to be found. The discovery rate is calibrated on endpoint counts, not discovery curves. Where hard decisions multiply with task length, reliability stays hard.",
    "method": "Theory plus synthesis: two propositions and a corollary about intervention libraries, with the mode-discovery rate calibrated on endpoint counts from three published failure taxonomies (ErrorAtlas, a HumanEval error categorisation, MWPES-300K), and the framework checked against about 60 published results, including 28 intervention citations across six capability axes and re-audits of five steep-decay papers.",
    "summary_html": [
      "Error analysis tends to merge three questions that the paper keeps apart: where errors occur, what recurs, and what fixes them. Errors occur at a sparse set of key decisions, and only a fraction of those, the <em>hard</em> ones, actually fail. What recurs is a local catalogue of failure modes. The fixes come from a still smaller library of capability interventions. The setting is a deployment <em>patch</em>: task inputs, schemas, users, retrieval corpus, evaluator, policy constraints and workflow horizon, all held fixed over a time window. Two modelling choices carry the math. Coverage by the top <em>m</em> interventions follows a log-head curve, <code>ln m / ln |C|</code>, used as a planning approximation; and the number of distinct modes found after <em>T</em> observed hard failures is posited, not derived, to grow at most logarithmically in <em>T</em>, with the rate calibrated on endpoint counts from ErrorAtlas, a HumanEval error categorisation and MWPES-300K. On the evidence side the paper synthesises about 60 published results, among them 28 intervention citations across six capability axes and re-audits of five papers often cited for steep decay.",
      "Proposition 1 is the negative half. If a domain keeps producing failures that escape any finite set of fixes for the earlier ones, its catalogue is infinite and no finite intervention dictionary can hold every mode below a fixed tolerance. What it rules out is a worst-case, mode-by-mode guarantee, not low average error; an uncovered tail can still carry little probability mass. Proposition 2 is the patch-local positive half: <code>m ≥ ⌈|C_eff|^(1 − ε/e_hard)⌉</code> interventions suffice to bring the per-hard-decision error rate from <code>e_hard</code> down to <code>ε</code>. If active modes grow logarithmically with hard decisions, and key decisions grow logarithmically with length, the budget grows doubly-logarithmically in sequence length; once the patch catalogue saturates, it stops depending on length. Read backwards, the discovery bound gives Corollary 1. At the conservative calibration σ ≈ 1.85, and if discovery runs close to the bound, five more modes cost about 15× more observed failures and ten about 220×. The published record fits. Taxonomies report small, head-heavy catalogues, targeted interventions largely close their own cluster while residuals land in other classes, and each re-audited steep-decay result decays over task structure, such as compositional graph size, fact count or evidence scope, rather than raw token length."
    ],
    "key_numbers": [
      {
        "label": "HumanEval failures from two error types",
        "value": "86.35%",
        "context": "AssertionError plus NameError, across 14 LLMs (Wen et al., 2024)"
      },
      {
        "label": "Failure categories in ErrorAtlas",
        "value": "17",
        "context": "head-concentrated categories from 83 models on 35 datasets and on the order of 10^4 failures (Ashury-Tahan et al., 2026); the anchor for σ ≈ 1.85"
      },
      {
        "label": "Extra failures needed to find five more modes",
        "value": "≈15×",
        "context": "inverse discovery cost at σ ≈ 1.85, assuming discovery tracks the logarithmic bound; ten more modes need ≈220×",
        "bad": true
      },
      {
        "label": "GSM-Hard score with PAL",
        "value": "20.1% → 61.5%",
        "context": "arithmetic offloaded to Python execution (Gao et al., 2023a); the remaining errors sit in comprehension, a different cluster"
      },
      {
        "label": "Planning prior for a patch library",
        "value": "≈ 50",
        "context": "named interventions that cover the bulk of per-hard-token failure mass in many measured domains; a prior to refine per patch, not a constant"
      }
    ],
    "editorial_html": [
      "For a team running a model inside one job, contract extraction say, or code repair, the unit of reliability work moves from the model to the patch. The paper's planning advice is to budget for a library of tens of interventions, then refine it from the patch's own discovery curve and failure ranking. Its estimate that about 50 named interventions cover the bulk of per-hard-token failure mass in many measured domains is a prior, not a constant; the same base model in cardiology RAG, legal drafting and code review needs three different libraries. The library is also coarser than the error list. One Python interpreter removes the execution part of arithmetic, unit conversion, counting, list manipulation and date arithmetic, which is why the paper counts in capability axes, six of them, rather than in named clusters.",
      "The metric matters as much as the budget. The polylogarithmic result is per hard decision, which suits systems that correct continuously. A one-shot system that needs the whole output right faces a stricter target, and as that target tightens the required library approaches full-catalogue coverage; service-level targets should follow the cost structure a system actually has. The levers sit in the context scaffold: what is retrieved, what stays in memory, which tools and validators are attached, how multi-turn state is orchestrated. Those choices decide which failure modes a patch can reach and which are covered.",
      "This is Part 2 of the error trilogy. <em>Beyond Exponential Decay</em> located long-context reliability at a handful of key decision points; this paper asks what goes wrong at those points, and argues that inside a patch it repeats. It closes by naming the engineering object without saying how the deployment scaffold should govern the library over time. <em>Frontier and Localhost</em> takes that up. It starts from the prompts, rules, memories, tools and eval suites that production teams already change outside the weights, and formalises a disciplined alternative to maintaining them as patchwork, which it calls artifact-layer descent."
    ],
    "limitations": "The logarithmic discovery rate is a postulate, not a theorem. Its calibration, σ between 0.87 and 1.85, rests on endpoint category counts from three taxonomies (general, code and math) rather than on discovery curves, and the paper notes that no subsample-discovery curve had been published for any LLM failure taxonomy at the time of writing; that measurement is the test the framework names. If discovery follows a Heaps power law instead, the doubly-logarithmic rate fails, though the budget stays polylogarithmic. The log-head coverage curve is a planning approximation, additivity across interventions is approximate, σ ≈ 1.85 is a single-point calibration, the hard fraction β is latent, and the gap between researcher-labelled categories and the latent modes they approximate is unmeasured. The calibration is untested on agentic workflows, long scientific reasoning and multi-turn tool use over millions of tokens, where the catalogue may grow faster than logarithmically, and a library calibrated on one patch under-covers the next. Most fundamentally, the framework relocates long-context difficulty rather than removing it: where the number of hard decisions grows with task length, reliability stays hard.",
    "concept_terms": [
      "patch-local reliability",
      "failure-mode catalogue",
      "intervention dictionary",
      "logarithmic mode discovery",
      "inverse discovery cost",
      "cluster-selective interventions"
    ],
    "references": [
      "Arbuzov et al. (2025). Beyond exponential decay: Rethinking error accumulation in large language models.",
      "Ashury-Tahan et al. (2026). ErrorMap and ErrorAtlas: Charting the failure landscape of large language models.",
      "Wen et al. (2024). Fixing function-level code generation errors for foundation large language models.",
      "Sun et al. (2025). Error classification of large language models on math word problems: A dynamically adaptive framework.",
      "Gao et al. (2023a). PAL: Program-aided language models."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/architecture-of-errors.txt",
    "verification": {
      "claims": [
        {
          "claim": "No finite fix list covers open-ended LLM use (claim_title, tldr, meta_description, gist)",
          "verdict": "CONFIRMED",
          "source_quote": "A fixed list of interventions cannot cover open-ended LLM use"
        },
        {
          "claim": "Inside a deployment patch the sufficient intervention list levels off once the catalogue saturates (claim_title, tldr)",
          "verdict": "CONFIRMED",
          "source_quote": "becomes domain-constant once the patch catalogue saturates"
        },
        {
          "claim": "Inside a patch, empirical evidence suggests failures recur in a small catalogue (tldr)",
          "verdict": "CONFIRMED",
          "source_quote": "Within such patches, empirical evidence suggests failures are sparse, repetitive, and concentrated in a small recurring catalogue"
        },
        {
          "claim": "The usual worry: if every token can fail, the chance of a fully correct long output collapses (gist)",
          "verdict": "CONFIRMED",
          "source_quote": "the chance of a fully correct n-token output is ( 1 − e ) n , collapsing"
        },
        {
          "claim": "The object to track is a catalogue of recurring failure modes, not sequence length (gist)",
          "verdict": "CONFIRMED",
          "source_quote": "We argue the object to track is not raw sequence length but a small, local catalogue of recurring failure modes."
        },
        {
          "claim": "Across open-ended use, failure modes can appear without bound (gist)",
          "verdict": "CONFIRMED",
          "source_quote": "new intervention-distinguishable failure modes can appear without bound"
        },
        {
          "claim": "Proposition 1: if each new failure escapes any finite dictionary covering the earlier ones, no finite dictionary holds every mode below a fixed tolerance (gist, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "no finite intervention dictionary can guarantee residual error below ε for every intervention-distinguishable mode in D"
        },
        {
          "claim": "Proposition 1 rules out a worst-case, mode-by-mode guarantee, not low average error; an uncovered tail can carry little mass (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "This is a worst-case, mode-wise guarantee, not a claim about expected residual error: under a distributional metric an infinite uncovered tail may still carry arbitrarily small mass."
        },
        {
          "claim": "Deployed systems run inside bounded patches such as legal review, code repair or contract extraction, with recurring tasks, schemas and evaluators (gist, editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "They operate inside operationally bounded patches (legal review, medical RAG, code repair, customer-support agents, contract extraction) with recurring tasks, schemas, tools, and evaluator expectations"
        },
        {
          "claim": "Published taxonomies report small, head-heavy catalogues (gist, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "(typically 8–20) with high top-mode coverage"
        },
        {
          "claim": "Proposition 2: under log active-mode exposure and head-heavy coverage, a sufficient per-hard-decision budget grows polylogarithmically in sequence length and stops growing once the patch catalogue saturates (gist)",
          "verdict": "CONFIRMED",
          "source_quote": "a sufficient per-hard-decision intervention budget grows polylogarithmically in sequence length and becomes domain-constant once the patch catalogue saturates"
        },
        {
          "claim": "Corollary 1: the minimum number of observed hard failures grows exponentially with the number of modes to be found (gist, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "accommodating q distinct discovered modes requires at least T ≥ exp (( q − A D ) /σ D ) observed hard failures"
        },
        {
          "claim": "The discovery rate is calibrated on endpoint counts, not discovery curves (gist, limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "These are endpoint counts, not discovery curves"
        },
        {
          "claim": "Where hard decisions multiply with task length, reliability stays hard (gist, limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "where the number of hard decisions itself grows with task length, reliability remains hard"
        },
        {
          "claim": "Theory: two propositions and one corollary (method)",
          "verdict": "CONFIRMED",
          "source_quote": "We formalize this transition with two propositions and one corollary."
        },
        {
          "claim": "The three calibration taxonomies are ErrorAtlas, a HumanEval error categorisation and MWPES-300K (method, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "ErrorAtlas (Ashury-Tahan et al., 2026), MWPES-300K (Sun et al., 2025), the HumanEval categorisation (Wen et al., 2024)"
        },
        {
          "claim": "The paper synthesises about 60 published results (method, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "It synthesises ∼ 60 published results"
        },
        {
          "claim": "28 intervention citations across six capability axes (method, summary, editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "A dedicated harvest yields 28 capability-elimination citations across six independent axes"
        },
        {
          "claim": "Re-audits of five papers often cited for steep decay (method, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "Five prominent papers are routinely cited as evidence that LLM reliability decays steeply with length"
        },
        {
          "claim": "Three questions kept apart: where errors occur, what recurs, what fixes them; fixes come from a smaller library (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "where errors occur (a sparse subset of key decisions), what recurs (a local catalogue of failure modes), and what fixes them (a smaller library of capability interventions)"
        },
        {
          "claim": "Only a fraction of key decisions, the hard ones, carry the actual failures (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "only a fraction concentrate the actual failures"
        },
        {
          "claim": "A patch fixes task inputs, schemas, users, retrieval corpus, evaluator, policy constraints and workflow horizon over a time window (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "A deployment patch D is not a topic label but an operational tuple fixing, over a time window, which failure modes are reachable"
        },
        {
          "claim": "Top-m coverage follows the log-head curve ln m / ln |C|, used as a planning approximation (summary, limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "as a planning approximation to F emp , not the true distribution"
        },
        {
          "claim": "Logarithmic mode discovery is posited as an empirical postulate, not derived (summary, limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "We therefore state logarithmic mode discovery as an empirical postulate, defensible by direct measurement, not a theorem."
        },
        {
          "claim": "Proposition 2: m ≥ ⌈|C_eff|^(1 − ε/e_hard)⌉ interventions suffice to bring the per-hard-decision error from e_hard to ε (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "Taking the ceiling (since m is integer-valued) yields m ≥ ⌈| C eff | 1 − ε/e hard ⌉"
        },
        {
          "claim": "If active modes grow logarithmically with hard decisions and key decisions grow logarithmically with length, the budget grows doubly-logarithmically in sequence length (summary; reworded from the elliptical 'and key decisions with length')",
          "verdict": "CONFIRMED",
          "source_quote": "If C active,D ( n ) grows logarithmically with the number of hard decisions and k ( n ) = Θ ( log n ) , then the pre-cap sufficient budget grows doubly-logarithmically in sequence length."
        },
        {
          "claim": "Once the patch catalogue saturates, the budget stops depending on length (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "which no longer depends on n"
        },
        {
          "claim": "At σ ≈ 1.85, under tightness, five more modes cost about 15× more observed failures and ten about 220× (summary, key number ≈15×; recomputed exp(5/1.85) = 14.9, exp(10/1.85) = 222.6)",
          "verdict": "CONFIRMED",
          "source_quote": "under tightness at σ D ≈ 1.85, five extra modes cost ≈ 15 × more observed failures and ten cost ≈ 220 ×"
        },
        {
          "claim": "σ ≈ 1.85 is the conservative calibration (summary, key number)",
          "verdict": "CONFIRMED",
          "source_quote": "we carry σ ≈ 1.85 as a conservative planning value"
        },
        {
          "claim": "Targeted interventions largely close their own cluster while residuals land in other classes (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "each intervention is cluster-selective (residuals land in a different cluster)"
        },
        {
          "claim": "Each re-audited steep-decay result decays over task structure such as compositional graph size, fact count or evidence scope, not raw token length (summary)",
          "verdict": "CONFIRMED",
          "source_quote": "showing it decays over task-structure variables — compositional graph size, fact count, log-time horizon, capacity threshold, evidence scope — rather than raw token length"
        },
        {
          "claim": "Key number 86.35%: AssertionError plus NameError across 14 LLMs on HumanEval (Wen et al., 2024); 63.64 + 22.71 = 86.35",
          "verdict": "CONFIRMED",
          "source_quote": "Reports AssertionError 63.64% + NameError 22.71% = 86.35% of HumanEval failures across 14 LLMs."
        },
        {
          "claim": "Key number 17: ErrorAtlas head-concentrated categories from 83 models on 35 datasets and about 10^4 failures (17 / ln 10^4 = 1.85)",
          "verdict": "CONFIRMED",
          "source_quote": "83 models × 35 datasets, ≳ 10 4 failures, 17 head-concentrated categories."
        },
        {
          "claim": "ErrorAtlas is the anchor for σ ≈ 1.85 (key number context)",
          "verdict": "CONFIRMED",
          "source_quote": "Conservative crossdomain anchor (used as planning value)"
        },
        {
          "claim": "Key number 20.1% → 61.5%: PAL on GSM-Hard, arithmetic offloaded to Python, residuals in comprehension (Gao et al., 2023a)",
          "verdict": "CONFIRMED",
          "source_quote": "PAL (Gao et al., 2023a) lifts GSM-Hard 20.1% → 61.5% (residuals in comprehension)"
        },
        {
          "claim": "Key number ≈ 50: named interventions covering the bulk of per-hard-token failure mass in many measured domains; a planning prior, not a constant (key number, editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "≈ 50 named interventions cover the bulk of the per-hard-token failure mass in many measured domains. This is a planning prior,"
        },
        {
          "claim": "Reliability work moves from the model to the patch (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "Reliability engineering is local patch coverage."
        },
        {
          "claim": "Budget for a library of tens of interventions, then refine from the patch's discovery curve and failure ranking (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "A team budgets initially for a library of tens of interventions, then refines from the local discovery curve"
        },
        {
          "claim": "The same base model in cardiology RAG, legal drafting and code review needs three different libraries (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "the same base model in cardiology RAG, legal drafting, and code review yields three different libraries"
        },
        {
          "claim": "One Python interpreter removes the execution part of arithmetic, unit conversion, counting, list manipulation and date arithmetic (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "a single Python interpreter removes the execution-error component of arithmetic, unit conversion, simple counting, list manipulation, and date arithmetic"
        },
        {
          "claim": "The library is coarser than the error list; the six capability axes are the better accounting unit (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "the six axes of Appendix B are the better accounting unit"
        },
        {
          "claim": "The polylog result is per hard decision, suited to continuously correcting systems; one-shot sequence-level targets are stricter and SLAs should follow the actual cost structure (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "per-hard-token residual error (continuous-correction systems) has a polylog budget, but sequence-level failure probability (one-shot systems) is strictly stricter"
        },
        {
          "claim": "As the one-shot target tightens, the required library approaches full-catalogue coverage (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "the required library approaches full-catalogue coverage"
        },
        {
          "claim": "Context-scaffold levers (retrieval, memory, tools and validators, multi-turn orchestration) decide which failure modes are reachable and covered (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "determine which failure modes in C D are reachable and which are covered"
        },
        {
          "claim": "Beyond Exponential Decay located long-context reliability at a handful of key decision points (editorial; checked against abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "abstracts.yaml: long-context reliability hinges on a handful of decision points, not on uniform per-token accuracy"
        },
        {
          "claim": "This paper is the next step after Beyond Exponential Decay, Part 2 of the error trilogy (editorial; series position from content/papers.yaml part: 2, sequence supported by the paper text)",
          "verdict": "CONFIRMED",
          "source_quote": "This paper takes the next step."
        },
        {
          "claim": "This paper asks what goes wrong at the key decision points and argues that inside a patch it repeats (editorial; corrected from 'finds', since patch-local recurrence rests on suggestive evidence and a modelling assumption)",
          "verdict": "CONFIRMED",
          "source_quote": "Sparsity tells us where errors live; the follow-up is what they are."
        },
        {
          "claim": "The paper names the engineering object but not how the deployment scaffold should govern the library over time (editorial)",
          "verdict": "CONFIRMED",
          "source_quote": "The paper names the engineering object (a patch-local failure catalogue and the budget covering its head) but not how the deployment-time scaffold"
        },
        {
          "claim": "Frontier and Localhost starts from the prompts, rules, memories, tools and eval suites production teams change outside the weights (editorial; checked against abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "abstracts.yaml: After deployment failures, teams modify prompts, rules, memories, skills, tools, eval suites"
        },
        {
          "claim": "Frontier and Localhost formalises a disciplined alternative to patchwork maintenance, called artifact-layer descent (editorial; checked against abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "abstracts.yaml: This paper formalises a disciplined alternative, which we call artifact-layer descent."
        },
        {
          "claim": "Calibration σ between 0.87 and 1.85 (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "σ ∈ [ 0.87, 1.85 ]"
        },
        {
          "claim": "The calibration rests on three taxonomies: general, code and math (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "The empirical anchor rests on three 2025–2026 taxonomies (general/code/math)"
        },
        {
          "claim": "No subsample-discovery curve had been published for any LLM failure taxonomy at the time of writing; that measurement is the named test (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "no subsample-discovery curve has been published for any LLM failure-mode taxonomy at this writing"
        },
        {
          "claim": "Heaps power-law discovery would break the doubly-logarithmic rate but keep the budget polylogarithmic (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "genuinely Heaps-power-law discovery would invalidate the doubly-logarithmic special case while preserving the qualitative polylog conclusion"
        },
        {
          "claim": "Additivity across interventions is approximate (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "inter-cluster additivity is only approximate"
        },
        {
          "claim": "σ ≈ 1.85 is a single-point calibration, β is latent, the label-to-latent-mode gap is unmeasured (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "the σ ≈ 1.85 estimate is a single-point calibration; β is latent; and the L2–L3 granularity gap is unmeasured"
        },
        {
          "claim": "Untested on agentic workflows, long scientific reasoning and multi-turn tool use over millions of tokens, where the catalogue may grow faster than logarithmically (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "untested on agentic workflows, long scientific reasoning, and multi-turn tool use over millions of tokens — where | C | may grow faster than logarithmically"
        },
        {
          "claim": "A library calibrated on one patch under-covers the next (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "so a library calibrated on one patch under-covers the next"
        },
        {
          "claim": "The framework relocates long-context difficulty rather than removing it (limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "the framework relocates long-context difficulty rather than resolving it"
        },
        {
          "claim": "Reference present: Arbuzov et al. (2025), Beyond exponential decay",
          "verdict": "CONFIRMED",
          "source_quote": "Beyond exponential decay: Rethinking error accumulation in large language models"
        },
        {
          "claim": "Reference present: Ashury-Tahan et al. (2026), ErrorMap and ErrorAtlas",
          "verdict": "CONFIRMED",
          "source_quote": "ErrorMap and ErrorAtlas: Charting the failure landscape of large language models"
        },
        {
          "claim": "Reference present: Wen et al. (2024), Fixing function-level code generation errors",
          "verdict": "CONFIRMED",
          "source_quote": "Fixing functionlevel code generation errors for foundation large language models"
        },
        {
          "claim": "Reference present: Sun et al. (2025), Error classification of LLMs on math word problems",
          "verdict": "CONFIRMED",
          "source_quote": "Error classification of large language models on math word problems: A dynamically adaptive framework"
        },
        {
          "claim": "Reference present: Gao et al. (2023a), PAL",
          "verdict": "CONFIRMED",
          "source_quote": "PAL: Program-aided language models"
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}