---
type: paper
slug: frontier-and-localhost
title: 'Frontier and Localhost: How Production AI Learns Outside the Weights'
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Ziwei Dong
- Dmitri Kalaev
- Alexey Shvets
date: '2026-06-23'
status: COLM 2026 workshop submission
line: Error-accumulation
pages: 13
html: https://telegrapher.ai/research/frontier-and-localhost/
pdf: https://telegrapher.ai/papers/frontier-and-localhost/frontier-and-localhost.pdf
reader: https://telegrapher.ai/research/frontier-and-localhost/read/
json: https://telegrapher.ai/api/papers/frontier-and-localhost.json
---

# Frontier and Localhost: How Production AI Learns Outside the Weights

## Paper gist

- **Claim:** Of ~130 LLM systems, none pairs automated scaffold fixes with governed cross-org promotion
- **TL;DR:** Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations.
- **Method:** A survey of roughly 130 LLM systems published or productionised between January 2023 and May 2026, scored on a uniform 0–5 rubric across six scaffold substrates and tagged by evidence tier, read against a formal model of gated scaffold updates whose descent inequality and convergence proposition are stated in an appendix.
- **Key result:** LLM systems surveyed: ~130; Systems in the core corpus: ~90; Systems passing the composite two-loop audit: thirteen
- **Why it matters:** For teams running agents in production, the practical change is small and specific: treat each rule, memory entry, skill or tool binding as an update with a measured effect, a rollback path and a scope, not as a note left for the next engineer.
- **Limits:** The evidence is a survey and a formal argument.
- **Status:** COLM 2026 workshop submission, June 2026
- **Read:** reader /research/frontier-and-localhost/read/, PDF /papers/frontier-and-localhost/frontier-and-localhost.pdf

## Abstract

Production LLM systems increasingly adapt outside the model weights. After deployment failures, teams modify prompts, rules, memories, skills, tools, eval suites, routing graphs, and governance pipelines on a cadence that frontier weight updates cannot match. But this scaffold layer is still mostly maintained as patchwork: fixes are proposed by intuition, committed with weak credit assignment, accumulated without pruning, and promoted beyond the scope where they were validated. This paper formalises a disciplined alternative, which we call artifact-layer descent. The pipeline is concrete. Recurring residuals are assigned to scaffold coordinates. Candidate artifact deltas are tested against patch loss. Accepted deltas persist with rollback. Promotion across contexts is bounded by evidence radius: a local fix spreads only as far as the evidence supports. Surveying approximately 130 production and research systems from 2023 to 2026, we find current systems partially instantiate this loop but leave a central architecture gap, namely governed scaffold optimization across contexts. The contribution is not the claim that prompts are weights. It is to name the missing optimizer for an adaptation layer the field has already built.

## What we did and found

The paper treats the layer around a frozen model as the thing being fit. It splits that scaffold into six substrates (instructions, skills, memory, tools, orchestration, governance), each an independently editable coordinate, and defines patch loss as the deployment's failure rate on its own recurring tasks. The disciplined loop, artifact-layer descent, starts from the accumulated failure record rather than from single incidents. It clusters failures into recurring modes and takes them in order of frequency times severity. For each mode it traces the coordinate responsible, such as a stale memory entry, a missing tool argument or a misrouted agent edge, and synthesises one candidate edit, which a gate keeps only if measured loss falls by a margin. A second loop decides how far an accepted edit travels. There the gate tightens as the evidence radius widens, because an edit that lowers a narrow average can raise a broad one. Against this model, roughly 130 systems from January 2023 to May 2026 are scored 0–5 per substrate; production claims are not pooled with peer-reviewed evidence.

About 90 systems clear the bar for patch-local adaptation, and roughly 38 more that were described as self-improving, scaffolded or agentic but fall short of it are kept as a contrast set. Every mature system has a recognisable version of each substrate. None matures all six. Research systems learn fast and lack enterprise governance; production systems govern well and learn slowly; open standards such as AGENTS.md and MCP make artifacts portable without making them adaptive. A stricter two-loop audit is passed by thirteen research systems, and when the surveyed two-loop systems are indexed by the kind of gate on each loop, they fill three cells of a 3 × 3 matrix. The missing architecture is defined by four properties: an automated gate on local edits, promotion between organisations, versioned lineage with role-based access, and rollback of a bad promotion without a redeploy. The closest systems split these between them. No surveyed system holds all four, and none holds the cross-organisational three under any kind of gate.

## Key numbers

| Measure | Value |
|---|---|
| LLM systems surveyed | ~130 |
| Systems in the core corpus | ~90 |
| Systems passing the composite two-loop audit | thirteen |
| SkillOpt's average margin over the strongest baseline | +5.4 points |
| Gain from LLM-authored skills over no skills | +0.0 pp |

## Why it matters

For teams running agents in production, the practical change is small and specific: treat each rule, memory entry, skill or tool binding as an update with a measured effect, a rollback path and a scope, not as a note left for the next engineer. Much of the machinery already exists under other names. An eval that blocks a merge is a validation check; a PR review against a shared rule repository is a promotion gate. What tends to be missing sits between the two, in tracing a failure back to the coordinate that caused it and in pruning edits that no longer earn their place.

The framing also says where patchwork breaks. Instruction files bloat. Persistent memory can be poisoned by a single malicious entry that is read back long after the attack, tool outputs carry prompt injection across vendors, and the eval suite that gates edits drifts along with the scaffold it is supposed to judge. The paper treats each of these as a missing piece of the optimizer, not as hygiene for later.

This is Part 3 of the error trilogy. Beyond Exponential Decay argues that long-context reliability hinges on a handful of decision points. The Architecture of Errors argues that inside an operationally bounded patch, failures fall into a small recurring catalogue, so reliability becomes a matter of discovering that catalogue and covering it. Frontier and Localhost follows the covering into production, where it happens in the scaffold. A residual made of a few repeating modes, each attachable to a scaffold coordinate, is what makes a local optimizer tractable at all; it also tells the optimizer where to spend.

## What this does not show

The evidence is a survey and a formal argument. Systems were found by searching public discourse and graded by evidence tier, from peer-reviewed papers down to industry-blog documentation. The two-loop count moves under two natural relaxations of the audit criterion, though the qualitative findings hold under both. The full corpus, the contrast set and the gate-family mapping are deferred to an extended version, so readers get a representative spot-check rather than every scored system. The descent reading is mechanism where its five conditions hold and analogy elsewhere. Where it applies, the convergence proposition promises a gate-stable scaffold rather than a global optimum, and in a deployment whose task mix shifts it tracks a moving optimum instead of settling. Whether the missing cross-organisational architecture should be built at all depends on the deployment; in safety-critical domains a reviewer's judgement in the gate may be a design feature rather than a defect. And the gates are themselves part of the moving scaffold, so keeping them held-out and recalibrated remains an open problem.

## Blog post

[Production AI learns outside the weights, without the optimizer that layer needs](https://telegrapher.ai/blog/frontier-and-localhost.md)

## Cite

```bibtex
@misc{arbuzov2026frontier,
  title         = {Frontier and Localhost: How Production AI Learns Outside the Weights},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},
  year          = {2026},
  note          = {COLM 2026 workshop submission},
  url           = {https://telegrapher.ai/research/frontier-and-localhost/}
}
```
