---
type: paper
slug: architecture-of-errors
title: 'The Architecture of Errors: From Universal Impossibility to Patch-Local LLM
  Reliability'
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Ziwei Dong
- Dmitri Kalaev
- Alexey Shvets
date: '2026-08-04'
status: COLM 2026 workshop poster
line: Error-accumulation
pages: 24
html: https://telegrapher.ai/research/architecture-of-errors/
pdf: https://telegrapher.ai/papers/architecture-of-errors/architecture-of-errors.pdf
reader: https://telegrapher.ai/research/architecture-of-errors/read/
json: https://telegrapher.ai/api/papers/architecture-of-errors.json
arxiv: https://arxiv.org/abs/2605.30628
openreview: https://openreview.net/forum?id=DimJCP8Bz1
---

# The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability

## Paper gist

- **Claim:** No finite fix list covers open-ended LLM use; inside a deployment, a sufficient list levels off
- **TL;DR:** No finite fix list covers every failure mode of open-ended LLM use. Inside one deployment, published taxonomies suggest failures recur in a small catalogue, so a sufficient fix library grows slowly with sequence length, then levels off.
- **Method:** Theory plus synthesis: two propositions and a corollary about intervention libraries, with the mode-discovery rate calibrated on endpoint counts from three published failure taxonomies (ErrorAtlas, a HumanEval error categorisation, MWPES-300K), and the framework checked against about 60 published results, including 28 intervention citations across six capability axes and re-audits of five steep-decay papers.
- **Key result:** HumanEval failures from two error types: 86.35%; Failure categories in ErrorAtlas: 17; Extra failures needed to find five more modes: ≈15×
- **Why it matters:** For a team running a model inside one job, contract extraction say, or code repair, the unit of reliability work moves from the model to the patch.
- **Limits:** The logarithmic discovery rate is a postulate, not a theorem.
- **Status:** COLM 2026 workshop poster, August 2026
- **Read:** reader /research/architecture-of-errors/read/, PDF /papers/architecture-of-errors/architecture-of-errors.pdf, arXiv:2605.30628 https://arxiv.org/abs/2605.30628

## Abstract

Reliability is the implicit goal of context management — what a model retrieves, remembers, and is scaffolded with is chosen so it fails less — yet it is usually analysed asymptotically, as if each token compounded the risk. We argue the object to track is not raw sequence length but a small, local catalogue of recurring failure modes. Universal LLM reliability is not a finite-library problem: across all possible tasks, tools, schemas, knowledge sources, and evaluator expectations, new intervention-distinguishable failure modes can appear without bound, so no finite intervention dictionary can guarantee bounded residual error for every such mode. But deployed systems do not operate over the whole universe. They operate inside operationally bounded patches (legal review, medical RAG, code repair, customer-support agents, contract extraction) with recurring tasks, schemas, tools, and evaluator expectations—the operational envelope that the context scaffold of retrieval, memory, tools, and orchestration defines and maintains. Within such patches, empirical evidence suggests failures are sparse, repetitive, and concentrated in a small recurring catalogue, so reliability becomes a local catalogue-discovery and intervention-coverage problem rather than an exponential token-length problem. We formalize this transition with two propositions and one corollary. Proposition 1 is the worst-case-mode-wise negative result: no finite intervention dictionary covers every distinguishable failure mode of an unbounded domain. Corollary 1 is the inverse-discovery implication: the logarithmic upper bound on mode discovery cannot accommodate linearly more distinct tail modes without exponentially more observed hard-failure events. Proposition 2 is the positive patch-local result: under log active-mode exposure and headheavy coverage, a sufficient per-hard-decision intervention budget grows polylogarithmically in sequence length and becomes domain-constant once the patch catalogue saturates. The framework relocates rather than dissolves long-context difficulty: where the number of hard decisions itself grows with task length, reliability remains hard; the contribution is to identify the on-axis intervention rather than to make those regimes easy.

## What we did and found

Error analysis tends to merge three questions that the paper keeps apart: where errors occur, what recurs, and what fixes them. Errors occur at a sparse set of key decisions, and only a fraction of those, the hard ones, actually fail. What recurs is a local catalogue of failure modes. The fixes come from a still smaller library of capability interventions. The setting is a deployment patch: task inputs, schemas, users, retrieval corpus, evaluator, policy constraints and workflow horizon, all held fixed over a time window. Two modelling choices carry the math. Coverage by the top m interventions follows a log-head curve, ln m / ln |C|, used as a planning approximation; and the number of distinct modes found after T observed hard failures is posited, not derived, to grow at most logarithmically in T, with the rate calibrated on endpoint counts from ErrorAtlas, a HumanEval error categorisation and MWPES-300K. On the evidence side the paper synthesises about 60 published results, among them 28 intervention citations across six capability axes and re-audits of five papers often cited for steep decay.

Proposition 1 is the negative half. If a domain keeps producing failures that escape any finite set of fixes for the earlier ones, its catalogue is infinite and no finite intervention dictionary can hold every mode below a fixed tolerance. What it rules out is a worst-case, mode-by-mode guarantee, not low average error; an uncovered tail can still carry little probability mass. Proposition 2 is the patch-local positive half: m ≥ ⌈|C_eff|^(1 − ε/e_hard)⌉ interventions suffice to bring the per-hard-decision error rate from e_hard down to ε. If active modes grow logarithmically with hard decisions, and key decisions grow logarithmically with length, the budget grows doubly-logarithmically in sequence length; once the patch catalogue saturates, it stops depending on length. Read backwards, the discovery bound gives Corollary 1. At the conservative calibration σ ≈ 1.85, and if discovery runs close to the bound, five more modes cost about 15× more observed failures and ten about 220×. The published record fits. Taxonomies report small, head-heavy catalogues, targeted interventions largely close their own cluster while residuals land in other classes, and each re-audited steep-decay result decays over task structure, such as compositional graph size, fact count or evidence scope, rather than raw token length.

## Key numbers

| Measure | Value |
|---|---|
| HumanEval failures from two error types | 86.35% |
| Failure categories in ErrorAtlas | 17 |
| Extra failures needed to find five more modes | ≈15× |
| GSM-Hard score with PAL | 20.1% → 61.5% |
| Planning prior for a patch library | ≈ 50 |

## Why it matters

For a team running a model inside one job, contract extraction say, or code repair, the unit of reliability work moves from the model to the patch. The paper's planning advice is to budget for a library of tens of interventions, then refine it from the patch's own discovery curve and failure ranking. Its estimate that about 50 named interventions cover the bulk of per-hard-token failure mass in many measured domains is a prior, not a constant; the same base model in cardiology RAG, legal drafting and code review needs three different libraries. The library is also coarser than the error list. One Python interpreter removes the execution part of arithmetic, unit conversion, counting, list manipulation and date arithmetic, which is why the paper counts in capability axes, six of them, rather than in named clusters.

The metric matters as much as the budget. The polylogarithmic result is per hard decision, which suits systems that correct continuously. A one-shot system that needs the whole output right faces a stricter target, and as that target tightens the required library approaches full-catalogue coverage; service-level targets should follow the cost structure a system actually has. The levers sit in the context scaffold: what is retrieved, what stays in memory, which tools and validators are attached, how multi-turn state is orchestrated. Those choices decide which failure modes a patch can reach and which are covered.

This is Part 2 of the error trilogy. Beyond Exponential Decay located long-context reliability at a handful of key decision points; this paper asks what goes wrong at those points, and argues that inside a patch it repeats. It closes by naming the engineering object without saying how the deployment scaffold should govern the library over time. Frontier and Localhost takes that up. It starts from the prompts, rules, memories, tools and eval suites that production teams already change outside the weights, and formalises a disciplined alternative to maintaining them as patchwork, which it calls artifact-layer descent.

## What this does not show

The logarithmic discovery rate is a postulate, not a theorem. Its calibration, σ between 0.87 and 1.85, rests on endpoint category counts from three taxonomies (general, code and math) rather than on discovery curves, and the paper notes that no subsample-discovery curve had been published for any LLM failure taxonomy at the time of writing; that measurement is the test the framework names. If discovery follows a Heaps power law instead, the doubly-logarithmic rate fails, though the budget stays polylogarithmic. The log-head coverage curve is a planning approximation, additivity across interventions is approximate, σ ≈ 1.85 is a single-point calibration, the hard fraction β is latent, and the gap between researcher-labelled categories and the latent modes they approximate is unmeasured. The calibration is untested on agentic workflows, long scientific reasoning and multi-turn tool use over millions of tokens, where the catalogue may grow faster than logarithmically, and a library calibrated on one patch under-covers the next. Most fundamentally, the framework relocates long-context difficulty rather than removing it: where the number of hard decisions grows with task length, reliability stays hard.

## Blog post

[No finite fix list covers every LLM failure, and a deployment doesn't need one](https://telegrapher.ai/blog/architecture-of-errors.md)

## Cite

```bibtex
@misc{arbuzov2026the,
  title         = {The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
  year          = {2026},
  note          = {COLM 2026 workshop poster},
  eprint        = {2605.30628},
  archivePrefix = {arXiv},
  url           = {https://telegrapher.ai/research/architecture-of-errors/}
}
```
