telegrapher

No finite fix list covers every LLM failure, and a deployment doesn't need one

The Architecture of Errors: From Universal Impossibility to Patch-Local LLM ReliabilityThe paper: summary, reader and PDF

Pool the HumanEval failures of more than a dozen LLMs, sort them by error type, and two types, AssertionError and NameError, account for 86.35% of the pile. ErrorAtlas, assembled from dozens of models on dozens of datasets, files its failures into 17 categories, with the mass bunched near the top. Taxonomies of math errors, multi-hop QA, tool-using agents and RAG come out the same shape: short lists with heavy heads.

The usual reliability analysis counts something else. Give every token an error probability e, and an n-token output survives with probability (1 − e)^n, which heads to zero. In the paper we follow the failure logs to a different object to manage: not output length but a catalogue of recurring failure modes. In Beyond Exponential Decay we had argued that a small share of tokens, 5–10%, are key decisions that depend on long-range context. That located the errors. The logs say what they are.

No finite list of fixes covers open-ended use

Keep a list of fixes: an interpreter for arithmetic, a constrained decoder for output schemas, a retriever for facts. A new task brings a new tool and fails in a way nothing on the list handles. You add a fix, and the next task does it again. If the domain is open enough that each new failure escapes any finite set of fixes for the earlier ones, the catalogue is infinite, and no finite list can hold every mode below a chosen error tolerance.

That is Proposition 1. It blocks the misreading that some fixed list of about 50 patterns covers LLM use in general, and it is narrower than it sounds. What it rules out is a worst-case guarantee, mode by mode. An infinite uncovered tail can still carry very little probability, so the result says nothing about average error.

Inside a deployment, failures tend to repeat

Deployed systems do not face open-ended use. They run in what we call a patch, a deployment that holds its working conditions fixed over a time window: task inputs, schemas, users, the retrieval corpus, the evaluator, policy constraints and the workflow horizon. Legal review under one jurisdiction is a patch; mix jurisdictions and it shifts. Several of those coordinates are exactly what a context-management system controls.

Inside a patch the question becomes how large the local library must be to cover enough of the catalogue. The framework keeps three layers apart. Errors occur at a sparse set of key decisions. A fraction of those, the hard decisions, actually fail, and what recurs is the catalogue of ways they fail. The fixes form a smaller library still.

A patch’s fix budget grows slowly, then stops

The natural plan is to rank a patch’s failure modes by their share of the hard-error mass and fix from the top. We model the share the top m fixes cover as ln m / ln |C|, with |C| the catalogue size, as a planning approximation. On ErrorAtlas’s 17 categories, five fixes then cover a little over half the mass and ten about four fifths. So under this curve, five fixes are enough to halve the hard-error rate.

Proposition 2 generalises the calculation. To bring the per-hard-decision error rate from e_hard down to a target ε, a library of

m ≥ ⌈|C|^(1 − ε/e_hard)⌉

interventions is sufficient. Halving makes the exponent one half, and the square root of 17, rounded up, is five. Per sequence, if active modes grow logarithmically with hard decisions and key decisions grow logarithmically with length, the budget grows doubly-logarithmically in length. Per patch, once every reachable mode has turned up, it stops depending on length. The budget is sufficient under the approximation, though not proven minimal.

The tail is expensive to find

We posit that after T observed hard failures in a patch, the number of distinct modes found grows at most like A + σ ln T. Read backwards, that postulate prices the tail: finding q modes takes at least exp((q − A)/σ) failures, so each extra mode multiplies the bill. At σ ≈ 1.85, the conservative calibrated value, and if discovery runs close to the bound, five more modes take about 15× more observed failures and ten about 220×. That is Corollary 1.

The cost sits on novelty. Failures in known modes stay cheap to observe; an unseen mode is what gets expensive. A catalogue’s head is cheap and heavy, its tail costly and thin. If the postulate holds, that is why generic post-training shows diminishing reliability returns once a patch’s head modes are covered.

Fixes empty their own bins

The evidence is a synthesis of about 60 published results. Taxonomies typically name between 8 and 20 modes, and targeted interventions largely close their own cluster:

Failure cluster Intervention Reported change
Arithmetic (GSM-Hard) PAL: write Python, run it 20.1% → 61.5%
Code logic and types Execution feedback (Reflexion + AgentCoder) 80% → 96.3% pass@1
Reasoning steps Process supervision (Math-Shepherd) 28.6% → 43.5% with reranking
Format and schema Constrained decoding invalid-token probability zero by construction

On GSM8K, Program-of-Thoughts takes calculation errors out of the failure log altogether; what remains is almost all reasoning and misunderstanding. The interpreter emptied its own bin.

Across the 28 intervention citations we gathered, on six capability axes, the pattern comes in two strengths. Seven of the citations empty their target class by construction, as a constrained decoder does. The rest shrink their class, sharply or moderately, and the failures left over land in structurally different classes, which is why fixes on different layers roughly add up.

Thirteen of ErrorAtlas’s 17 categories are addressed by one of the six axes. The other four are failures of choosing what to do rather than of doing it, and they set the floor for residual error.

A library of tens, sized per patch

For a team running one patch, the framework turns into a budget. Start with tens of interventions, then measure: collect the patch’s hard failures, watch how fast new modes appear, rank them by mass, cover from the top. The paper’s planning prior, which it presents as consistent with the head-mass figures in ErrorAtlas, HumanEval and MWPES, is about 50 named interventions for the bulk of per-hard-token failure mass in many measured domains. Proposition 1 rules out that same number as a fixed list for LLM use in general. Inside one patch it is a starting budget, revised against the patch’s own failures: the same base model in cardiology RAG, legal drafting and code review needs three different libraries.

The library is also coarser than the error list. One Python interpreter takes out the execution part of arithmetic, unit conversion, counting, list manipulation and date arithmetic.

Mind the metric, though. The budget is per hard decision, which suits systems that correct as they go. A one-shot system that needs the whole output right faces a stricter target, and as that target tightens the library approaches full-catalogue coverage.

The curve that would settle it

The logarithmic discovery rate is a postulate, not a theorem. Its calibration, σ between 0.87 and 1.85, comes from endpoint counts in three taxonomies (general, code and math): how many categories a taxonomer assigned at one corpus size, not how that number grew. If discovery follows Heaps’ law instead, the doubly-logarithmic rate goes, though the budget stays polylogarithmic.

Other pieces are softer than the algebra makes them look. The finite patch catalogue is a modelling assumption: the taxonomies pool failures across models and benchmarks, and none of the paper’s evidence measures one deployment’s catalogue directly. The coverage curve is an approximation, and fixes add up only roughly; two that share a prompt channel can interfere. The hard fraction is latent, and how well labelled categories track the underlying modes is unmeasured. The calibration is untested on agentic workflows, long scientific reasoning and multi-turn tool use over millions of tokens, where the catalogue may grow faster, and a library calibrated on one patch under-covers the next.

Most fundamentally, the framework moves long-context difficulty rather than removing it. We re-audited five papers often cited for steep decay, and each decays over something other than raw token length, such as compositional graph size, supporting-fact count or evidence scope. Where hard decisions multiply with the task like that, reliability stays hard. The framework names the axis to intervene on. It does not make those regimes easy.

The paper also leaves governance open: it names the library, not how the scaffold around a model should manage it over time. Frontier and Localhost takes that up. Production teams already change prompts, rules, memories and tools outside the weights, mostly as patchwork, and that paper formalises a disciplined alternative it calls artifact-layer descent.

The test is easy to state. Take one patch, subsample its failures at growing sizes, and plot distinct modes against failures observed. When we wrote the paper, no such curve had been published for any LLM failure taxonomy. If it flattens, as the small endpoint counts hint without being able to show, the reliability question for a deployed system stops being whether n-token error can be bounded. It becomes whether enough of the patch’s failure modes have been catalogued.

Read the paperThe Architecture of Errors: From Universal Impossibility to Patch-Local LLM ReliabilityCOLM 2026 workshop poster, August 2026