Of ~130 LLM systems, none pairs automated scaffold fixes with governed cross-org promotion
Frontier and Localhost: How Production AI Learns Outside the Weights
Mikhail L Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, Alexey Shvets
COLM 2026 workshop submission, June 2026
Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations.
What we did and found
The paper treats the layer around a frozen model as the thing being fit. It splits that scaffold into six substrates (instructions, skills, memory, tools, orchestration, governance), each an independently editable coordinate, and defines patch loss as the deployment's failure rate on its own recurring tasks. The disciplined loop, artifact-layer descent, starts from the accumulated failure record rather than from single incidents. It clusters failures into recurring modes and takes them in order of frequency times severity. For each mode it traces the coordinate responsible, such as a stale memory entry, a missing tool argument or a misrouted agent edge, and synthesises one candidate edit, which a gate keeps only if measured loss falls by a margin. A second loop decides how far an accepted edit travels. There the gate tightens as the evidence radius widens, because an edit that lowers a narrow average can raise a broad one. Against this model, roughly 130 systems from January 2023 to May 2026 are scored 0–5 per substrate; production claims are not pooled with peer-reviewed evidence.
About 90 systems clear the bar for patch-local adaptation, and roughly 38 more that were described as self-improving, scaffolded or agentic but fall short of it are kept as a contrast set. Every mature system has a recognisable version of each substrate. None matures all six. Research systems learn fast and lack enterprise governance; production systems govern well and learn slowly; open standards such as AGENTS.md and MCP make artifacts portable without making them adaptive. A stricter two-loop audit is passed by thirteen research systems, and when the surveyed two-loop systems are indexed by the kind of gate on each loop, they fill three cells of a 3 × 3 matrix. The missing architecture is defined by four properties: an automated gate on local edits, promotion between organisations, versioned lineage with role-based access, and rollback of a bad promotion without a redeploy. The closest systems split these between them. No surveyed system holds all four, and none holds the cross-organisational three under any kind of gate.
Key numbers
| LLM systems surveyedproduction and research systems published or productionised between January 2023 and May 2026 | ~130 |
| Systems in the core corpusscore ≥ 3 on at least one of six substrates; roughly 38 more fail that bar and are kept as a contrast set | ~90 |
| Systems passing the composite two-loop auditresearch systems whose Loop 1 edits a persisted artifact from session signal and whose Loop 2 promotes it under an explicit gate; the count moves under looser criteria | thirteen |
| SkillOpt's average margin over the strongest baselineacross 52 (model, benchmark, harness) cells, as reported by Yang et al. (2026) for a single skill document optimised as trainable external state | +5.4 points |
| Gain from LLM-authored skills over no skillsagainst +16.2 pp for human-curated skills, as reported in work the paper cites; X. Zhang et al. (2026) attribute the gap to lifecycle management | +0.0 pp |
What this does not show
The evidence is a survey and a formal argument. Systems were found by searching public discourse and graded by evidence tier, from peer-reviewed papers down to industry-blog documentation. The two-loop count moves under two natural relaxations of the audit criterion, though the qualitative findings hold under both. The full corpus, the contrast set and the gate-family mapping are deferred to an extended version, so readers get a representative spot-check rather than every scored system. The descent reading is mechanism where its five conditions hold and analogy elsewhere. Where it applies, the convergence proposition promises a gate-stable scaffold rather than a global optimum, and in a deployment whose task mix shifts it tracks a moving optimum instead of settling. Whether the missing cross-organisational architecture should be built at all depends on the deployment; in safety-critical domains a reviewer's judgement in the gate may be a design feature rather than a defect. And the gates are themselves part of the moving scaffold, so keeping them held-out and recalibrated remains an open problem.
The authors' abstract
Production LLM systems increasingly adapt outside the model weights. After deployment failures, teams modify prompts, rules, memories, skills, tools, eval suites, routing graphs, and governance pipelines on a cadence that frontier weight updates cannot match. But this scaffold layer is still mostly maintained as patchwork: fixes are proposed by intuition, committed with weak credit assignment, accumulated without pruning, and promoted beyond the scope where they were validated. This paper formalises a disciplined alternative, which we call artifact-layer descent. The pipeline is concrete. Recurring residuals are assigned to scaffold coordinates. Candidate artifact deltas are tested against patch loss. Accepted deltas persist with rollback. Promotion across contexts is bounded by evidence radius: a local fix spreads only as far as the evidence supports. Surveying approximately 130 production and research systems from 2023 to 2026, we find current systems partially instantiate this loop but leave a central architecture gap, namely governed scaffold optimization across contexts. The contribution is not the claim that prompts are weights. It is to name the missing optimizer for an adaptation layer the field has already built.
Figures



The rest of the trilogy
Cite
@misc{arbuzov2026frontier,
title = {Frontier and Localhost: How Production AI Learns Outside the Weights},
author = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
year = {2026},
note = {COLM 2026 workshop submission},
url = {https://telegrapher.ai/research/frontier-and-localhost/}
}Builds on
- Tiwari et al. (2026). Learning, Fast and Slow: Towards LLMs That Adapt Continually.
- Kirk et al. (2024). Understanding the Effects of RLHF on LLM Generalisation and Diversity.
- Yang et al. (2026). SkillOpt: Executive Strategy for Self-Evolving Agent Skills.
- Agrawal et al. (2025). GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.
- Singh et al. (2025). The Leaderboard Illusion.