No finite fix list covers open-ended LLM use; inside a deployment, a sufficient list levels off
The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets
COLM 2026 workshop poster, August 2026. arXiv:2605.30628
No finite fix list covers every failure mode of open-ended LLM use. Inside one deployment, published taxonomies suggest failures recur in a small catalogue, so a sufficient fix library grows slowly with sequence length, then levels off.
What we did and found
Error analysis tends to merge three questions that the paper keeps apart: where errors occur, what recurs, and what fixes them. Errors occur at a sparse set of key decisions, and only a fraction of those, the hard ones, actually fail. What recurs is a local catalogue of failure modes. The fixes come from a still smaller library of capability interventions. The setting is a deployment patch: task inputs, schemas, users, retrieval corpus, evaluator, policy constraints and workflow horizon, all held fixed over a time window. Two modelling choices carry the math. Coverage by the top m interventions follows a log-head curve, ln m / ln |C|, used as a planning approximation; and the number of distinct modes found after T observed hard failures is posited, not derived, to grow at most logarithmically in T, with the rate calibrated on endpoint counts from ErrorAtlas, a HumanEval error categorisation and MWPES-300K. On the evidence side the paper synthesises about 60 published results, among them 28 intervention citations across six capability axes and re-audits of five papers often cited for steep decay.
Proposition 1 is the negative half. If a domain keeps producing failures that escape any finite set of fixes for the earlier ones, its catalogue is infinite and no finite intervention dictionary can hold every mode below a fixed tolerance. What it rules out is a worst-case, mode-by-mode guarantee, not low average error; an uncovered tail can still carry little probability mass. Proposition 2 is the patch-local positive half: m ≥ ⌈|C_eff|^(1 − ε/e_hard)⌉ interventions suffice to bring the per-hard-decision error rate from e_hard down to ε. If active modes grow logarithmically with hard decisions, and key decisions grow logarithmically with length, the budget grows doubly-logarithmically in sequence length; once the patch catalogue saturates, it stops depending on length. Read backwards, the discovery bound gives Corollary 1. At the conservative calibration σ ≈ 1.85, and if discovery runs close to the bound, five more modes cost about 15× more observed failures and ten about 220×. The published record fits. Taxonomies report small, head-heavy catalogues, targeted interventions largely close their own cluster while residuals land in other classes, and each re-audited steep-decay result decays over task structure, such as compositional graph size, fact count or evidence scope, rather than raw token length.
Key numbers
| HumanEval failures from two error typesAssertionError plus NameError, across 14 LLMs (Wen et al., 2024) | 86.35% |
| Failure categories in ErrorAtlashead-concentrated categories from 83 models on 35 datasets and on the order of 10^4 failures (Ashury-Tahan et al., 2026); the anchor for σ ≈ 1.85 | 17 |
| Extra failures needed to find five more modesinverse discovery cost at σ ≈ 1.85, assuming discovery tracks the logarithmic bound; ten more modes need ≈220× | ≈15× |
| GSM-Hard score with PALarithmetic offloaded to Python execution (Gao et al., 2023a); the remaining errors sit in comprehension, a different cluster | 20.1% → 61.5% |
| Planning prior for a patch librarynamed interventions that cover the bulk of per-hard-token failure mass in many measured domains; a prior to refine per patch, not a constant | ≈ 50 |
What this does not show
The logarithmic discovery rate is a postulate, not a theorem. Its calibration, σ between 0.87 and 1.85, rests on endpoint category counts from three taxonomies (general, code and math) rather than on discovery curves, and the paper notes that no subsample-discovery curve had been published for any LLM failure taxonomy at the time of writing; that measurement is the test the framework names. If discovery follows a Heaps power law instead, the doubly-logarithmic rate fails, though the budget stays polylogarithmic. The log-head coverage curve is a planning approximation, additivity across interventions is approximate, σ ≈ 1.85 is a single-point calibration, the hard fraction β is latent, and the gap between researcher-labelled categories and the latent modes they approximate is unmeasured. The calibration is untested on agentic workflows, long scientific reasoning and multi-turn tool use over millions of tokens, where the catalogue may grow faster than logarithmically, and a library calibrated on one patch under-covers the next. Most fundamentally, the framework relocates long-context difficulty rather than removing it: where the number of hard decisions grows with task length, reliability stays hard.
The authors' abstract
Reliability is the implicit goal of context management — what a model retrieves, remembers, and is scaffolded with is chosen so it fails less — yet it is usually analysed asymptotically, as if each token compounded the risk. We argue the object to track is not raw sequence length but a small, local catalogue of recurring failure modes. Universal LLM reliability is not a finite-library problem: across all possible tasks, tools, schemas, knowledge sources, and evaluator expectations, new intervention-distinguishable failure modes can appear without bound, so no finite intervention dictionary can guarantee bounded residual error for every such mode. But deployed systems do not operate over the whole universe. They operate inside operationally bounded patches (legal review, medical RAG, code repair, customer-support agents, contract extraction) with recurring tasks, schemas, tools, and evaluator expectations—the operational envelope that the context scaffold of retrieval, memory, tools, and orchestration defines and maintains. Within such patches, empirical evidence suggests failures are sparse, repetitive, and concentrated in a small recurring catalogue, so reliability becomes a local catalogue-discovery and intervention-coverage problem rather than an exponential token-length problem. We formalize this transition with two propositions and one corollary. Proposition 1 is the worst-case-mode-wise negative result: no finite intervention dictionary covers every distinguishable failure mode of an unbounded domain. Corollary 1 is the inverse-discovery implication: the logarithmic upper bound on mode discovery cannot accommodate linearly more distinct tail modes without exponentially more observed hard-failure events. Proposition 2 is the positive patch-local result: under log active-mode exposure and headheavy coverage, a sufficient per-hard-decision intervention budget grows polylogarithmically in sequence length and becomes domain-constant once the patch catalogue saturates. The framework relocates rather than dissolves long-context difficulty: where the number of hard decisions itself grows with task length, reliability remains hard; the contribution is to identify the on-axis intervention rather than to make those regimes easy.
Figure

The rest of the trilogy
Cite
@misc{arbuzov2026the,
title = {The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability},
author = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
year = {2026},
note = {COLM 2026 workshop poster},
eprint = {2605.30628},
archivePrefix = {arXiv},
url = {https://telegrapher.ai/research/architecture-of-errors/}
}Builds on
- Arbuzov et al. (2025). Beyond exponential decay: Rethinking error accumulation in large language models.
- Ashury-Tahan et al. (2026). ErrorMap and ErrorAtlas: Charting the failure landscape of large language models.
- Wen et al. (2024). Fixing function-level code generation errors for foundation large language models.
- Sun et al. (2025). Error classification of large language models on math word problems: A dynamically adaptive framework.
- Gao et al. (2023a). PAL: Program-aided language models.