telegrapher

Production AI learns outside the weights, without the optimizer that layer needs

Frontier and Localhost: How Production AI Learns Outside the WeightsThe paper: summary, reader and PDF

Say an agent in production keeps calling a booking tool without a time-zone argument. Someone adds a line to the instruction file telling it to pass one; the failures stop, so the line stays. Weeks later another team copies the file into a deployment whose tool takes no such argument, and the rule now sits where nobody tested it. The file keeps swelling. Nobody can say which line fixed what.

None of that touched the model. Frontier weights move a release at a time, months apart; the systems around them change continuously. In the paper we argue that a growing share of deployment-time adaptation now happens in this scaffold of instruction files, memories, skills, tools and evals, and that it is run mostly as patchwork: fixes proposed by intuition, kept without knowing what they did, left unpruned, and promoted past the scope where they were checked.

The weights are tuned for breadth, the scaffold for fit

To stay competent across unseen deployments, frontier training smooths over local conventions, contradictory tenant preferences and rare workflow-specific failures. Reliability inside one deployment is decided by exactly those details. Writing every deployment’s quirks into the weights fails on cost, release cadence and cross-deployment interference at once. A scaffold edit is cheap, stays local, can be reverted, and lands on a human timescale.

Familiar parts of the stack look different once you read them as one scaffold. An eval that blocks a merge is a validation check; a pull-request review against a shared rule repository is a promotion gate. None of these parts is new; the move is to read them together as one adaptation surface and ask what discipline it needs. (We are not claiming that prompts are weights.)

A fix stays only if the failure rate drops

The disciplined loop works from the accumulated failure record, not from single incidents. It clusters failures into recurring modes and ranks them by frequency times severity. For the top mode it finds the scaffold coordinate responsible, say a stale memory entry, and writes one candidate edit. Then it measures. A gate keeps the edit if the deployment’s failure rate on its own recurring tasks drops by a margin; a kept edit that later proves harmful is rolled back.

We call this artifact-layer descent, and that failure rate the patch loss. Patchwork becomes descent once five pieces are in place: a measurable patch loss, a record of failures, a failure traced back to the coordinate that caused it, a synthesised edit, and a gate that checks the edit lowers loss. Where all five hold, “descent” names the mechanism. Elsewhere it is an analogy.

Parts of the field already work this way. SkillOpt (Yang et al., 2026) treats a single markdown skill document as the trainable state of a frozen agent, keeps an edit only if it clears a held-out gate, and reports an average margin of 5.4 points over each setting’s strongest baseline. All five pieces are in place there, for one document.

Then comes the step most systems skip. Loop 1 fixes one deployment; Loop 2 decides whether to share the fix. Sharing changes the target, not the step size: an edit that lowers one deployment’s failure rate can raise a broader average. So a fix should travel only as far as its evidence reaches, a distance we call its evidence radius, and the gate tightens as that radius widens. Governance, on this reading, is the regularizer that keeps a local fix from being tuned against the wrong average.

Every mature system has six substrates; none matures all six

We surveyed roughly 130 systems from January 2023 to May 2026, found by searching for systems described as self-improving, scaffolded or agentic. About 90 keep a persistent artifact updated from deployment signal, scoring 3 or more on our 0–5 rubric for at least one substrate. The other 38 or so carry the name without clearing that bar; we kept them as a contrast set.

Substrate Systems (n) Mean score (0–5) Survey finding
Instructions 12 2.83 Rules persist; none closes both loops
Skills 16 2.62 Active in research; production governance weaker
Memory 19 3.16 Widespread; correctable, versioned memory is rare
Tools 19 3.63 Most mature in production
Orchestration 19 2.68 Topologies exist; gated evolution is rare
Governance 21 3.05 Eval-gated merge and rollback, typically no inner loop
Integrated 34 3.74 Highest mean, mostly research systems

The spread follows what each kind of system was built for. Research systems learn fast and lack enterprise governance; thirteen of them pass a stricter audit that asks for both loops at once. Production systems govern well and learn slowly. Open standards such as AGENTS.md and MCP have made artifacts portable across vendors, which is not the same as making them adapt.

The empty corner is cross-organisational

Four properties define the architecture we could not find: (a) an automated gate on local edits; (b) promotion that routes artifacts between organisations; (c) versioned lineage with role-based access; (d) rollback of a bad promotion without a redeploy. The nearest systems split them up. AlphaEvolve and SkillForge pair (a) with promotion inside one organisation. SkillOpt has (a) and part of (c). Governance-first platforms hold (c) and (d) but not (a) or (b).

No surveyed system combines all four. The cross-organisational three, (b) through (d), go unfilled under any kind of gate.

Patchwork fails in its own ways

Each failure mode the paper lists marks a missing piece of the optimizer. The main near-term risk is bloat: instruction-following degrades and latency climbs as the instruction count grows. Bloat also costs signal quality. In a comparison we cite, LLM-authored skills added +0.0 pp over a no-skill baseline against +16.2 pp for human-curated ones, a gap X. Zhang et al. (2026) attribute to lifecycle management. Pruning has to live inside the loop.

Because memory is read back into later prompts, one poisoned entry can be acted on long after the attack. And the eval suite that gates edits is itself a scaffold artifact; once its score becomes the target, it gets gamed.

What changes for a team running agents

Treat each rule, memory entry, skill or tool binding as an update with a measured effect, a rollback path and a scope, not as a comment left for the next engineer. The merge-blocking eval and the reviewed rule repository already do part of this. What tends to be missing sits between them: tracing a failure back to the coordinate that caused it, pruning edits that no longer earn their place, and sharing a fix only as far as its evidence reaches.

Where this sits in our work

This is the third paper in a three-part argument about errors. Beyond Exponential Decay argues that long-context reliability hinges on a handful of decision points; The Architecture of Errors, that inside an operationally bounded patch, failures are sparse, repetitive and concentrated in a small recurring catalogue. That is what makes a local optimizer workable here: a deployment’s residual is a short list of repeating modes, each attachable to a scaffold coordinate.

What is still open

The evidence is a survey and a formal argument; we did not build the missing architecture. Systems came from public sources, with evidence from peer-reviewed papers down to industry blogs, and the full corpus waits for an extended version. The two-loop count moves under two natural relaxations of the criterion, though the qualitative findings survive both.

The formal result is narrower than it may sound. Under stated assumptions the loss converges and the scaffold ends up gate-stable, with no admissible edit clearing the gate. That scaffold need not be a global optimum, and under a shifting task mix the loop tracks a moving optimum rather than settling. Whether the cross-organisational corner should be filled depends on the deployment; in safety-critical domains a reviewer’s judgement in the gate may be the design rather than the gap.

A closely related problem sits beside the missing architecture. The instruments that judge scaffold edits belong to the same moving scaffold: an eval suite adapts on the cadence of what it gates, so a reading taken before an edit goes stale as failures shift beneath it. The gates need the same discipline as the optimizer: held-out evals, recalibrated as the target moves. A gate is only as good as its coverage of the modes that recur, and those modes move.

Read the paperFrontier and Localhost: How Production AI Learns Outside the WeightsCOLM 2026 workshop submission, June 2026