# Frontier and Localhost: How Production AI Learns Outside the Weights

Full text, page by page. Paper page: https://telegrapher.ai/research/frontier-and-localhost.md

## Page 1

Under review as a conference paper at COLM 2026

Frontier and Localhost:
How Production AI Learns Outside the Weights

Anonymous authors
Paper under double-blind review

Abstract

Production LLM systems increasingly adapt outside the model weights.
After deployment failures, teams modify prompts, rules, memories, skills,
tools, eval suites, routing graphs, and governance pipelines on a cadence
that frontier weight updates cannot match. But this scaffold layer is still
mostly maintained as patchwork: fixes are proposed by intuition, committed with weak credit assignment, accumulated without pruning, and promoted beyond the scope where they were validated. This paper formalises
a disciplined alternative, which we call artifact-layer descent. The pipeline
is concrete. Recurring residuals are assigned to scaffold coordinates. Candidate artifact deltas are tested against patch loss. Accepted deltas persist
with rollback. Promotion across contexts is bounded by evidence radius: a
local fix spreads only as far as the evidence supports. Surveying approximately 130 production and research systems from 2023 to 2026, we find
current systems partially instantiate this loop but leave a central architecture gap, namely governed scaffold optimization across contexts. The
contribution is not the claim that prompts are weights. It is to name the
missing optimizer for an adaptation layer the field has already built.

1. The gradient moved outside the model

The standard mental model of an LLM system is the model: pre-training fixes the weights,
fine-tuning nudges them, and deployment wraps a thin prompt-and-tool layer around the
result and ships. Production in 2025 and 2026 has stopped behaving this way. Frontier
models update on cycles measured in months; the systems built on top of them update
continuously, in markdown instruction files, persistent memories, retrieved skills, registered
tools, evaluation suites, agent topologies, version-controlled prompt repositories, and PRgated governance. A growing share of deployment-time adaptation now occurs outside the
weights, in a fast-changing local scaffold wrapped around a slow-changing frontier. The
section title is a metaphor: the scaffold layer is not a gradient in the calculus sense. The
substantive claim is the milder one — production AI has accidentally created an external
learning surface, and right now it is being operated as patchwork rather than as a disciplined
optimization process.

Read together, these artifacts form a single external adaptation surface that current practice still treats as scattered engineering. A Claude Skill is a learned procedural weight kept
outside the model. A Cursor Rule is a local parameter for one project’s distribution shift. A
RAG connector is patch-specific memory access. A tool registry is capability provisioning
along an axis the base model cannot reach. A PR review against a shared rule repository is
a promotion gate on a federated update. An eval suite that blocks merge is a validation-loss
check before promotion. None of these mechanisms is new; what is new is reading them as
one adaptation surface and asking what optimization discipline that surface needs. This paper surveys roughly 130 production and research systems, identifies the six substrates where
the scaffold lives, formalises a two-loop architecture (local update; cross-context promotion
under a gate), and names a missing architecture as its most consequential open problem:
governed cross-tenant scaffold optimization with versioned lineage, role-based access, and
rollback. The single-line thesis is this: the field has built an external adaptation layer for
LLMs, but not the optimizer that should govern it.

1

## Page 2

Under review as a conference paper at COLM 2026

2. Why patch errors live outside the weights

The scaffold is not merely where production systems happen to store patches. Under
deployment-time constraints — release cadence, tenant isolation, inspectability, rollback
— it is often the practical substrate, because it is cheap, locally scoped, reversible, and
updatable at human time-scales while the weights are not.

Frontier training optimises for breadth. The weights are tuned for cross-patch generalisation, smoothing over local conventions, contradictory tenant preferences, and rare workflowspecific failures to preserve broad competence across millions of unseen deployments. RLHF
reduces output diversity and homogenises preferred styles across users (Kirk et al. 2024;
Padmakumar and He 2023), so the frontier model is trained to be unable to encode a
single tenant’s idiosyncrasies as preferred behaviour. Tiwari et al. (2026) supply the matching two-timescale argument from the in-context / in-weights side: in-context adaptation
is cheap and rapid but capacity-bounded, in-weights adaptation is expressive but induces
catastrophic forgetting and plasticity loss. Production reliability, by contrast, is achieved
in depth: inside one deployment patch, reliability is determined by exactly what global
training had to smooth away. Encoding every patch in the weights is ruled out on economic,
release-cycle, and cross-patch-interference grounds at once. The frontier model generalises;
the local scaffold specialises; reliability comes from governing that specialisation.

Prior work on error accumulation (Arbuzov et al. 2025, 2026) makes the specialisation
tractable. It argues that LLM errors concentrate sparsely at a few key tokens and cluster into
a finite catalogue of recurring failure modes whose size grows only slowly in the number of
observed failures. A patch’s residual is therefore not diffuse noise but a small set of repeating
modes, each attachable to a scaffold coordinate — which is what makes a disciplined local
optimizer possible at all, and what tells that optimizer where to spend its budget.

Three stages of scaffold maintenance recur across the corpus. They differ in how candidate
updates are produced, whether credit assignment is explicit, and whether promotion is
evidence-bounded. Most production deployments sit at Stage 1 or 2; the formalism of §3
applies to Stage 3.

Update
mechanism

Credit
assignment

Promotion
discipline

Dominant
failure mode

Implicit / ad
hoc
Weak; eval is
the only signal

None

Bloat

2. Random
artifact search

Intuition-driven
local edits
Many variants
tried under eval

Overfitting to
in-distribution
patch

3.
Artifact-layer
descent

Residual →
coordinate →
delta

Explicit per
failure mode

Eval-gated CI,
no failure-to-coordinate
tracing
Evidenceradius-bounded;
reversible

Stage

1. Patchwork

(the discipline
this paper
formalises)

A patchwork loop becomes artifact-layer descent once five pieces are in place: a measurable
patch loss, captured residuals, a failure traced back to the scaffold coordinate that caused
it, a synthesised candidate edit, and a gate that checks whether the edit actually lowers loss
before it is kept or promoted. The survey of §5 asks which deployed systems have closed
which of those five steps; §6 names the one closure no public system has yet reached.

3. Artifact-layer descent

The disciplined version of scaffold maintenance is one loop, run again and again over the life
of a deployment. It does not react to single failures; it works from their accumulated record.
As the deployment runs, it logs what went wrong — failed outputs, user corrections, rejected
answers, eval regressions, tool-call errors. Reflecting over that record, the loop clusters the

2

## Page 3

Under review as a conference paper at COLM 2026

failures into recurring failure modes and takes up the ones that matter most: a mode earns
attention in proportion to how often it recurs and how badly it fails — repeatability times
severity, which is exactly its contribution to the deployment’s failure rate. For such a mode it
locates the single scaffold coordinate responsible — a stale memory entry, an under-specified
instruction, a missing tool argument, a brittle retrieval rule, a misrouted agent edge — and
synthesises one candidate edit to it. It measures whether the edit lowers patch loss (the
deployment’s failure rate on its own recurring tasks), lets a gate accept or reject it on that
measurement, keeps an accepted edit, and prunes or rolls back one that a later check shows
made things worse. Then it takes the step most systems skip: it decides how far the edit
should travel, promoting it to other deployments only as far as its evidence reaches — the
edit’s evidence radius.

Most current production does not close this loop. A failure happens, someone adds a rule
or memory entry, the artifact pile grows, nobody knows what caused improvement or when
to prune, and fixes leak into broader contexts. That is patchwork, not descent. The loop
has two distinct knobs. Loop 1 is the local fix on one deployment’s scaffold. Loop 2 is the
separate decision of whether that fix should be shared more widely; it does not change the
step size but the target — which average failure rate the scaffold is being tuned against —
and the gate gets stricter as the radius widens, because an edit that lowers a narrow average
can raise a broad one. This is why governance is not administrative overhead: it is the
regularization system that keeps a small fix from being aimed at the wrong target.

Write the patch loss as L D ( θ s ) = E x ∼ D [ℓ( π ( θ M , θ s , x ))] , with the frontier weights θ M fixed
and the scaffold θ s — a structured object over the six substrate coordinates — the thing
being fit. A failure yields no analytic gradient, only a candidate edit z t proposed in order
of the recurrence-times-severity of the mode it addresses; the gate accepts it when its noisy
estimate of the directional loss change clears a margin,

b z L D ( θ s t ) + λ Ω ( z t ) ≤ − τ t ,
G t ( z t ) = 1 ⇐⇒ ∆
t

where Ω ≥ 0 penalises edit complexity and promotion breadth and τ t ≥ 0 is the required
margin. Production systems implement this without writing the equation: an eval blocks
a merge on regression, a reviewer rejects a vague rule, a governance system holds a local
artifact back until evidence accumulates. The descent reading is mechanism when the five
conditions above hold and analogy otherwise; Appendix A states the assumptions, a one-step
descent inequality, and a convergence proposition to a gate-stable scaffold.

Recent work operationalises all five conditions. SkillOpt (Yang et al. 2026) treats a single
markdown skill document as the trainable external state of a frozen agent and edits it with
explicit deep-learning-style controls — rollout and reflection batches, add/delete/replace
edits, a bounded textual learning rate with cosine decay, a held-out selection gate, and a
rejected-edit buffer kept as negative feedback — and is best or tied-best on all 52 evaluated
(model, benchmark, harness) cells, beating the strongest per-cell baseline by +5.4 points on
average and reporting positive cross-model, cross-harness, and cross-benchmark transfer that
reads as an empirical evidence-radius measurement rather than a governance mechanism.
SkillOpt is one instance in a fast-crystallising space: reflective prompt evolution (GEPA,
Agrawal et al. 2025), reflective context learning (Vassilyev et al. 2026), harness optimisation
(Lee et al. 2026; Lin et al. 2026; Cai et al. 2026), and lifecycle hygiene (X. Zhang et al.
2026) each instantiate a different subset of the loop, which confirms the architecture is real
rather than any one group’s local optimum.

4. Six scaffold substrates

A patch-local scaffold composes six substrates. They are the coordinates of θ s , each independently editable.

3

## Page 4

Under review as a conference paper at COLM 2026

Substrate

What persists

Score-5 mark

S1. Instructions

Project/team/org rules
attached to a workspace
Procedural artifacts
(code/text) loaded on
demand
Cross-session experience
(episodic / semantic /
procedural)
Vetted action surface +
context schemas

Versioned, two-loop,
Auto-gated promotion
Skill bank self-modifies on
failure, unit-test gated,
promoted
Provenance + versioning +
correction pathways

S2. Skills

S3. Memory

S4. Tools

S5. Orchestration

Multi-role graph + routing

S6. Governance

Versioning, promotion,
audit, rollback

Domain-curated bundle is
the product; eval-gated tool
versions
Population/policy updates
from outer-loop signal
dev→staging→prod with
eval thresholds + audit log
+ rollback

The 0–5 rubric is uniform: 0 ephemeral; 1 persistent local; 2 persistent with tool attachment but no feedback; 3 multi-scope or feedback without automated promotion; 4 automated
feedback updates the scaffold (one loop closed); 5 two-loop versioned promotion with governance (both loops closed). The descent argument is invariant to the exact partition so long
as coordinates remain independently editable.

5. Survey and findings

The survey covers roughly 130 LLM systems published or productionised between January
2023 and May 2026, assembled by searching public discourse for systems described as “selfimproving,” “scaffolded,” or “agentic,” then filtering by whether the system wraps a foundation model with at least one persistent artifact updated from deployment signal. About
90 satisfy this patch-local discriminator (score ≥ 3 on at least one substrate) and form the
core corpus; roughly 38 carry the name but fail it and are kept as a contrast corpus so the
boundary is auditable rather than rhetorical. Each system carries one or more evidence
tiers, from peer-reviewed publication (T1) down to industry-blog documentation (T5); production claims are not pooled with peer-reviewed evidence when computing cluster-level
statistics. A stricter composite-two-loop audit — Loop 1 modifies a persisted artifact from
session signal and Loop 2 has an explicit gated cross-context promotion mechanism — is
passed by thirteen research systems; the count moves under two natural relaxations of the
criterion, but the qualitative findings below survive both.

Substrate

n

Mean

Score-5
systems

Survey
finding

S1.
Instructions

12

2.83

(none)

Rules persist
across IDEs
and agents,
but none
closes both an
automated
inner loop
and a
governed
two-loop
promotion

4

## Page 5

Under review as a conference paper at COLM 2026

Substrate

n

Mean

S2. Skills

16

2.62

S3. Memory

19

3.16

S4. Tools

19

3.63

S5.
Orchestration

19

2.68

S6.
Governance

21

3.05

INTEGRATED

34

3.74

Score-5
systems

Survey
finding

SAGE,
COSPLAY

Research
frontier active
(failuretriggered skill
update);
production
governance
weaker
Cuadros et al. Cross-session
memory is
widespread;
correctable,
versioned
memory is
rare
MCP, Harvey, Most mature
Hippocratic
production
substrate;
tool bundles
are the
commercial
unit
(none)
Topologies
exist;
topology
evolution
under a gate
is rare
Braintrust,
Mature
Vellum,
prompts-as-
LangSmith
code:
eval-gated
merge, fast
rollback,
typically
without an
inner loop
NanoResearch, The
AutoAgent,
high-water
SkillRL
mark, but
dominated by
research
systems
lacking
enterprise
governance

Two things stand out. The field has independently converged on the six substrates — every
mature system has a recognisable instance of each, under different names — yet no system
matures all six at once: research systems learn fast and lack enterprise governance, production systems govern well and learn slowly, and open standards solve artifact portability
but not adaptation. The maturity gradient is not a quirk of the corpus; it tracks which
constraints each kind of system was built to satisfy.

5

## Page 6

Under review as a conference paper at COLM 2026

Four archetypal clusters recur (Appendix B). Research full-auto systems (NanoResearch,
SkillOpt, AlphaEvolve, DGM, Voyager) automate both loops under algorithmic gates and
uniformly lack enterprise governance. Production hybrid systems (SkillForge, CASCADE,
Sierra, Devin) pair an algorithmic Loop-1 gate with a reviewer-judgment Loop-2 gate — the
production-viable cell. Governance-first systems (Braintrust, Vellum, LangSmith Hub) automate Loop 2 alone via eval-gated CI, with no failure-triggered inner loop. Open standards
(AGENTS.md, MCP, Claude Skills, Cursor Rules) converge across vendors on artifact format and tool attachment; AGENTS.md alone spans 60{,}000+ repositories with measured
eﬀiciency gains (Lulla et al. 2026). Indexed by gate type on each loop, the composite systems populate only three cells of a 3 × 3 matrix — [Auto × Auto], [Auto × Human], and
[None × Auto] — leaving the human-authored-inner-loop cells empty for structural reasons.

6. The missing architecture: governed cross-tenant scaffold optimization

Four properties define an architecture that no surveyed system fully instantiates: (a) Autogated Loop 1, an algorithmic commit criterion; (b) multi-tenant cross-org promotion, Loop
2 routing artifacts between organisations; (c) versioned lineage with RBAC; and (d) rollback
of a bad promotion without redeploy. The closest approaches partition the requirements.
AlphaEvolve has (a) and partial (b) within one organisation; SkillForge has (a) and (b)
inside one enterprise with partial (c); SkillOpt has (a) and partial (c), and reports transfer
as evidence-radius measurement but implements neither (b) nor (d); CORAL (Qu et al.
2026) is the only asynchronous multi-agent Loop 2 in the corpus, within one system; PACE
(Ling et al. 2026) shows the two-timescale architecture under small-frozen-model constraints;
the governance-first cluster has (c) and (d) but neither (a) nor (b). No surveyed system
combines all four, and the cross-organisational triad (b, c, d) is itself unfilled under any gate
type. Naming the corner converts a vague gap into a four-property audit criterion any future
system can be measured against. Whether it should be filled is deployment-dependent —
in safety-critical domains a reviewer-judgment gate may be a design feature rather than a
defect — but no public system fills it today, and that is what matters for the next two years
of architecture work.

7. A new failure surface

Patchwork scaffold maintenance introduces failure modes weight-only systems do not have.
Each one is also a missing piece of the optimizer. Scaffold bloat is the dominant near-term
risk: instruction-following degrades and latency grows sharply as instruction count scales
(Jaroslawicz et al. 2025; Liu et al. 2023), and bloat is also a signal-quality problem —
LLM-authored skills have been reported at + 0.0 pp over a no-skill baseline against + 16.2
pp for human-curated ones, a gap (X. Zhang et al. 2026) attributes to lifecycle management
— so complexity penalties and explicit pruning must be first-class. Poisoned memory is the
dominant security risk: because memory is persistent and read back into later prompts, one
malicious entry can be retrieved and acted on long after the attack (Zou et al. 2024; Z.
Xu et al. 2026), which calls for provenance, write-gating, and a rollback envelope. Prompt
injection via tools and MCP makes a tool’s output a cross-vendor attack surface, so tool
bindings belong in the same promotion gate as any other coordinate. Coordination failures
in multi-agent topologies (Cemri et al. 2025) make orchestration a coordinate subject to the
same discipline. Patch overfitting is structural — a scaffold tuned to a deployment over-fits
it, every time — and is contained only by evidence-radius discipline at promotion. Plasticity
loss at the scaffold level (context rot, instruction conflict, memory pollution, tool-version
drift) makes expiry and rollback first-class operations rather than retrospective hygiene.

Eval fragility is the dominant ecosystem risk, and it compounds because the eval suite is
itself a patch-local artifact rather than a fixed instrument. It is optimised against: once
a metric becomes the target it is gamed, and access asymmetries alone can produce up
to 112% relative gains that reflect exposure to the test rather than capability (Singh et
al. 2025). It is read non-independently: the gate scores path-dependent, non-i.i.d. agent
trajectories, so its readings do not compose like ordinary test-set statistics, and evaluating
an optimizer only by its downstream gains cannot distinguish informed Loop-1 updates

6

## Page 7

Under review as a conference paper at COLM 2026

from trial-and-error (Ong et al. 2026). It goes stale: because the suite adapts on the same
cadence as the scaffold it gates, a reading taken before an edit ages as the patch’s own
failure distribution moves under it. And because failures concentrate in a finite catalogue
of recurring modes rather than spreading uniformly, a suite that samples tasks at random
under-weights exactly the families that drive the failure rate — the gate is only as good
as its coverage of those modes. The architectural response is held-out evals scoped within
evidence radius, weighted toward the recurring modes, re-read as the construct drifts, and
checked against an update’s attribution rather than only its score.

The failure surface is real, quantified by 2024–2026 work, and largely unmitigated in production. Every entry on it names a piece of the optimizer that current systems do not yet
have.

8. Conclusion

The field has built an external adaptation layer for LLMs, but not the optimizer that
should govern it. If reliability is increasingly repaired through persistent scaffold artifacts,
then production AI needs a theory and an engineering discipline for the discrete, gated,
auditable learning that happens there — not just for the differentiable learning inside the
model. This paper is one step toward that discipline: a survey of where production sits on
the patchwork → random-search → descent gradient, a formalisation of the Stage-3 limit
under stated assumptions, and a named missing architecture where the optimizer most
plainly does not yet exist.

The frontier model generalises; the local scaffold specialises; reliability comes from governing
that specialisation. The most consequential open architecture problem combines automated
local evolution with auditable cross-tenant aggregation: Auto-gated Loop 1 plus governed
multi-tenant Loop 2, with versioned lineage, RBAC, and rollback. A closely related problem
rides alongside it. The instruments that validate these edits are themselves part of the
moving, clustered scaffold, so the same evidence-radius and lifecycle discipline the optimizer
needs has to extend to the gates that judge it — held-out, family-aware, and recalibrated as
the target moves. Turning today’s patchwork into tomorrow’s optimizer depends on both.

References

Agashe, Saaket et al. 2025. “Agent S2: A Compositional Generalist-Specialist Framework
for Computer Use Agents.” arXiv Preprint. https://arxiv.org/abs/2504.00906.

Agashe, Saaket, Jiuzhou Han, Shuyu Gan, Jiachen Yang, Ang Li, and Xin Eric Wang. 2024.
“Agent S: An Open Agentic Framework That Uses Computers Like a Human.” arXiv
Preprint. https://arxiv.org/abs/2410.08164.

AGENTS.md. 2025. AGENTS.md: A Cross-Vendor Open Standard for Agent Instructions.
Linux Foundation. https://agents.md.

Agrawal, Lakshya A., Shangyin Tan, Dilara Soylu, et al. 2025. “GEPA: Reflective Prompt
Evolution Can Outperform Reinforcement Learning.” arXiv Preprint. https://arxiv.org/
abs/2507.19457.

Alzubi, Salaheddin et al. 2026. “EvoSkill: Automated Skill Discovery for Multi-Agent
Systems.” arXiv Preprint. https://arxiv.org/abs/2603.02766.

Anthropic. 2024. Model Context Protocol Specification. https://modelcontextprotocol.io.

Arbuzov, Mikhail L., Sisong Bei, Ziwei Dong, Dmitri Kalaev, and Alexey A. Shvets. 2025.
“Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models.”
arXiv Preprint. https://arxiv.org/abs/2505.24187.

7

## Page 8

Under review as a conference paper at COLM 2026

Arbuzov, Mikhail L., Alexey A. Shvets, and Sisong Bei. 2026. The Architecture of Errors: Logarithmic Mode Discovery and Polylogarithmic Intervention Budgets for Long-
Context LLM Reliability.

Cai, Qianshu, Yonggang Zhang, Xianzhang Jia, et al. 2026. “MOSS: Self-Evolution Through
Source-Level Rewriting in Autonomous Agent Systems.” arXiv Preprint. https://arxiv.
org/abs/2605.22794.

Cemri, Mert et al. 2025. “Why Do Multi-Agent LLM Systems Fail?” arXiv Preprint.
https://arxiv.org/abs/2503.13657.

Chen, Minghao, Yihang Li, Yanting Yang, Shiyu Yu, Binbin Lin, and Xiaofei He. 2024.
“AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning.” Advances in Neural Information Processing Systems (NeurIPS).
https://arxiv.org/abs/2405.16247.

Chen, Yixing et al. 2025. “Multi-Agent Evolve: LLM Self-Improve Through Co-Evolution.”
arXiv Preprint. https://arxiv.org/abs/2510.23595.

Cuadros, Diego F., Abdoul-Aziz Maiga, Helen Meskhidze, and Andre Curtis-Trudel. 2026.
“Governed Collaborative Memory as Artificial Selection in LLM-Based Multi-Agent Systems.” arXiv Preprint. https://arxiv.org/abs/2605.04264.

Dohare, Shibhansh et al. 2024. “Loss of Plasticity in Deep Continual Learning.” Nature
632 (8026): 768–74. https://doi.org/10.1038/s41586-024-07711-7.

Fawzi, Alhussein, Matej Balog, Aja Huang, et al. 2022. “Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning.” Nature 610 (7930): 47–53.

Harvey. 2026. Harvey Product Overview. https://www.harvey.ai.

He, Yufei et al. 2025. “EvoTest: Evolutionary Test-Time Learning for Self-Improving
Agentic Systems.” arXiv Preprint. https://arxiv.org/abs/2510.13220.

Hippocratic AI. 2026. Polaris Clinical Outcome Evidence. https://www.hippocraticai.com.

Hu, Shengran, Cong Lu, and Jeff Clune. 2024. “Automated Design of Agentic Systems
(ADAS).” arXiv Preprint. https://arxiv.org/abs/2408.08435.

Huang, Xu et al. 2025. “CASCADE: Cumulative Agentic Skill Creation Through Autonomous Development and Evolution.” arXiv Preprint. https://arxiv.org/abs/2512.
23880.

Jaroslawicz, Daniel et al. 2025. “How Many Instructions Can LLMs Follow at Once?” arXiv
Preprint. https://arxiv.org/abs/2507.11538.

Kirk, Robert, Ishita Mediratta, Christoforos Nalmpantis, et al. 2024. “Understanding the
Effects of RLHF on LLM Generalisation and Diversity.” ICLR 2024. https://arxiv.org/
abs/2310.06452.

Lee, Yoonho, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea
Finn. 2026. “Meta-Harness: End-to-End Optimization of Model Harnesses.” arXiv
Preprint. https://arxiv.org/abs/2603.28052.

Li, Hanchen, Runyuan He, Qizheng Zhang, et al. 2026. “Combee: Scaling Prompt Learning
for Self-Improving Language Model Agents.” arXiv Preprint. https://arxiv.org/abs/
2604.04247.

8

## Page 9

Under review as a conference paper at COLM 2026

Lin, Jiahang, Shichun Liu, Chengjun Pan, et al. 2026. “Agentic Harness Engineering:
Observability-Driven Automatic Evolution of Coding-Agent Harnesses.” arXiv Preprint.
https://arxiv.org/abs/2604.25850.

Ling, Chen, Pei Chen, Albert Guan, et al. 2026. “PACE: Two-Timescale Self-Evolution for
Small Language Model Agents.” arXiv Preprint. https://arxiv.org/abs/2605.23019.

Liu, Nelson F. et al. 2023. “Lost in the Middle: How Language Models Use Long Contexts.”
Transactions of the Association for Computational Linguistics (TACL). https://arxiv.
org/abs/2307.03172.

Liu, Xingyan et al. 2026. “SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills
in Cloud Technical Support.” arXiv Preprint. https://arxiv.org/abs/2604.08618.

Lulla, Jai Lal et al. 2026. “On the Impact of AGENTS.md Files on the Eﬀiciency of AI
Coding Agents.” arXiv Preprint. https://arxiv.org/abs/2601.20404.

Lyle, Clare, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will
Dabney. 2023. “Understanding Plasticity in Neural Networks.” International Conference
on Machine Learning (ICML). https://arxiv.org/abs/2303.01486.

Madaan, Aman et al. 2023. “Self-Refine: Iterative Refinement with Self-Feedback.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2303.
17651.

Ni, Jingwei et al. 2026. “Trace2Skill: Distill Trajectory-Local Lessons into Transferable
Agent Skills.” arXiv Preprint. https://arxiv.org/abs/2603.25158.

Novikov, Alexander et al. 2025. AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery. arXiv preprint arXiv:2506.13131. https://arxiv.org/abs/2506.13131.

Ong, Kai Tzu-iunn, Minseok Kang, Dongwook Choi, et al. 2026. “Towards Direct Evaluation of Harness Optimizers via Priority Ranking.” arXiv Preprint. https://arxiv.org/
abs/2605.22505.

Padmakumar, Vishakh, and He He. 2023. “Does Writing with Language Models Reduce
Content Diversity?” arXiv Preprint. https://arxiv.org/abs/2309.05196.

Park, Joon Sung et al. 2023. “Generative Agents: Interactive Simulacra of Human Behavior.” Proceedings of the 36th Annual ACM Symposium on User Interface Software and
Technology (UIST). https://arxiv.org/abs/2304.03442.

Qu, Ao, Han Zheng, Zijian Zhou, et al. 2026. “CORAL: Towards Autonomous Multi-
Agent Evolution for Open-Ended Discovery.” arXiv Preprint. https://arxiv.org/abs/
2604.01658.

Shalev, Yuval, Zifeng Ding, and Mateja Jamnik. 2026. “Training Language Agents to Learn
from Experience.” arXiv Preprint. https://arxiv.org/abs/2605.20477.

Shinn, Noah et al. 2023. “Reflexion: Language Agents with Verbal Reinforcement Learning.”
Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/
2303.11366.

Sierra. 2024. Sierra Agent OS. https://sierra.ai.

Sierra. 2025. Agent OS 2.0: From Answers to Memory and Action. https://sierra.ai/blog/
agent-os-2-0.

9

## Page 10

Under review as a conference paper at COLM 2026

Singh, Shivalika et al. 2025. “The Leaderboard Illusion.” arXiv Preprint. https://arxiv.
org/abs/2504.20879.

Song, Weijia, Jiashu Yue, and Zhe Pang. 2026. “ABSTRAL: Automatic Design of
Multi-Agent Systems Through Iterative Refinement and Topology Optimization.” arXiv
Preprint. https://arxiv.org/abs/2603.22791.

Sumers, Theodore R., Shunyu Yao, Karthik Narasimhan, and Thomas L. Griﬀiths. 2023.
“Cognitive Architectures for Language Agents (CoALA).” Transactions on Machine
Learning Research (TMLR). https://arxiv.org/abs/2309.02427.

Tan, Weihao et al. 2024. “Cradle: Empowering Foundation Agents Towards General Computer Control.” arXiv Preprint. https://arxiv.org/abs/2403.03186.

Tanjim, Md Mehrab, Jayakumar Subramanian, Xiang Chen, et al. 2026. “MOCHA: Multi-
Objective Chebyshev Annealing for Agent Skill Optimization.” arXiv Preprint. https:
//arxiv.org/abs/2605.19330.

Tiwari, Rishabh, Kusha Sareen, Lakshya A. Agrawal, et al. 2026. “Learning, Fast and Slow:
Towards LLMs That Adapt Continually.” arXiv Preprint. https://arxiv.org/abs/2605.
12484.

Tran, Dat, and Douwe Kiela. 2026. “Single-Agent LLMs Outperform Multi-Agent Systems
on Multi-Hop Reasoning Under Equal Thinking Token Budgets.” arXiv Preprint. https:
//arxiv.org/abs/2604.02460.

Vassilyev, Nikita, William Berrios, Ruowang Zhang, Bo Han, Douwe Kiela, and Shikib
Mehri. 2026. “Reflective Context Learning: Studying the Optimization Primitives of
Context Space.” arXiv Preprint. https://arxiv.org/abs/2604.03189.

Wang, Guanzhi et al. 2023. “Voyager: An Open-Ended Embodied Agent with Large
Language Models.” arXiv Preprint. https://arxiv.org/abs/2305.16291.

Wang, Jiongxiao et al. 2025. “Reinforcement Learning for Self-Improving Agent with Skill
Library.” arXiv Preprint. https://arxiv.org/abs/2512.17102.

Wang, Xiaoxing et al. 2026. “AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents.” arXiv Preprint. https://arxiv.org/abs/2603.09716.

Wang, Zhenhailong et al. 2025. “Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks.” arXiv Preprint. https://arxiv.org/abs/2501.11733.

Wu, Rong et al. 2025. “EvolveR: Self-Evolving LLM Agents Through an Experience-Driven
Lifecycle.” arXiv Preprint. https://arxiv.org/abs/2510.16079.

Wu, Xiyang et al. 2026. “Co-Evolving LLM Decision and Skill Bank Agents for Long-
Horizon Tasks.” arXiv Preprint. https://arxiv.org/abs/2604.20987.

Wu, Zhiyong et al. 2024. “OS-Copilot: Towards Generalist Computer Agents with Self-
Improvement.” arXiv Preprint. https://arxiv.org/abs/2402.07456.

Xia, Peng et al. 2026. “SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.” arXiv Preprint. https://arxiv.org/abs/2602.08234.

Xu, Jinhang et al. 2026. “NanoResearch: Co-Evolving Skills, Memory, and Policy for
Personalized Research Automation.” arXiv Preprint. https://arxiv.org/abs/2605.10813.

Xu, Zhenlin et al. 2026. “From Storage to Steering: Memory Control Flow Attacks on LLM
Agents.” arXiv Preprint. https://arxiv.org/abs/2603.15125.

10

## Page 11

Under review as a conference paper at COLM 2026

Yang, Yifan, Ziyang Gong, Weiquan Huang, et al. 2026. SkillOpt: Executive Strategy for
Self-Evolving Agent Skills. arXiv preprint arXiv:2605.23904. https://microsoft.github.
io/SkillOpt.

Yao, Shunyu, Jeffrey Zhao, Dian Yu, et al. 2023. “ReAct: Synergizing Reasoning and
Acting in Language Models.” International Conference on Learning Representations
(ICLR). https://arxiv.org/abs/2210.03629.

Yuksekgonul, Mert, Federico Bianchi, Joseph Boen, et al. 2024. “TextGrad: Automatic
‘Differentiation’ via Text.” arXiv Preprint. https://arxiv.org/abs/2406.07496.

Zehle, Tom. 2026. “CANTANTE: Optimizing Agentic Systems via Contrastive Credit
Attribution.” arXiv Preprint. https://arxiv.org/abs/2605.13295.

Zhang, Hanrong et al. 2026. “CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary
Verification.” arXiv Preprint. https://arxiv.org/abs/2604.01687.

Zhang, Jenny et al. 2025. “Darwin Gödel Machine: Open-Ended Evolution of Self-
Improving Agents.” arXiv Preprint. https://arxiv.org/abs/2505.22954.

Zhang, Qizheng, Changran Hu, Shubhangi Upasani, et al. 2025. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. arXiv preprint
arXiv:2510.04618. https://arxiv.org/abs/2510.04618.

Zhang, Xing, Yanwei Cui, Guanghui Wang, et al. 2026. “Ratchet: A Minimal Hygiene
Recipe for Self-Evolving LLM Agents.” arXiv Preprint. https://arxiv.org/abs/2605.
22148.

Zhou, Yingli, Wang Shu, Yaodong Su, et al. 2026. “A Comprehensive Survey on Agent
Skills: Taxonomy, Techniques, and Applications.” arXiv Preprint. https://arxiv.org/
abs/2605.07358.

Zou, Wei et al. 2024. “PoisonedRAG: Knowledge Corruption Attacks to Retrieval-
Augmented Generation of Large Language Models.” arXiv Preprint. https://arxiv.org/
abs/2402.07867.

Appendix A. Formalization of artifact-layer descent

This appendix supplies the verification machinery for §3: the setup, a one-step descent
inequality under a noisy estimator, a convergence proposition with its assumptions, and the
optimisation-analogy correspondence.

Setup. Let the frontier model have fixed parameters θ M and let the scaffold be a structured
object θ s ∈ Θ s = Θ 1 × · · · × Θ 6 over the six substrate coordinates; Θ s is discrete, symbolic,
and versioned. For a deployment patch D, L D ( θ s ) = E x ∼ D [ℓ( π ( θ M , θ s , x ))] . A candidate
edit z t ∼ Q t ( z | H t , θ s t ) is drawn from a proposal process conditioned on the accumulated
failure history H t . Q t is not uniform: for a mode m that recurs with frequency f m and carries
per-occurrence loss s m , its expected contribution is f m s m — repeatability times severity —
so high-frequency or high-severity modes are proposed first. The finite directional difference
b z L D .
is ∆ z L D ( θ s ) = L D ( θ s ⊕ z ) − L D ( θ s ) ; the gate sees only a noisy estimate ∆

b z L D |
Descent inequality. Assume (A1) positive descent bias on accepted edits — E [ ∆
t
G t = 1 ] ≤ − γ t − τ t for some γ t > 0 — and (A2) bounded estimator bias — | E [ ∆ z t L D −
b z L D | G t = 1 ]| ≤ ε t . With p t = Pr ( G t = 1 | θ s t ) and the gated update of §3,
∆
t

E [ L D ( θ s t + 1 ) | θ s t ] ≤ L D ( θ s t ) − p t ( γ t + τ t ) + p t ε t .

11

## Page 12

Under review as a conference paper at COLM 2026

Per-step expected improvement is the gate’s margin minus its estimator bias, weighted by
acceptance probability.

Convergence (informal). Add (A3) proposal coverage (at every non-gate-stable scaffold, Q t
places probability ≥ q > 0 on an admissible edit with E [ ∆ z L D ] ≤ − δ), (A4) effective
compactness of the reachable orbit, and (A5) persistent acceptance with summable bias
( ∑ t p t ( γ t + τ t ) = ∞, ∑ t p t ε t < ∞), with 0 ≤ L D ≤ B. Then L D ( θ s t ) is a non-negative
supermartingale up to summable noise, converges almost surely, and the limit set is gatestable: no admissible edit clears the gate with negative expected directional loss in the
limit. This is convergence of the loss and stability of the gate decision, not convergence
to a global optimum — a globally suboptimal gate-stable scaffold is consistent with the
result, exactly as for local stochastic optimization under noisy estimators. In non-stationary
patches (D = D t ) the same mechanism tracks a moving optimum rather than converging
once, which is why recency weighting, pruning, expiry, and rollback are the scaffold-level
analogues of continual-learning machinery rather than optional hygiene.

Optimization concept

Scaffold analogue

Parameter vector

Instructions, skills, memories, tools, routing
graphs, eval gates, governance rules
Failure, user correction, eval regression,
retrieval miss, tool error, schema violation
Credit assignment identifying the
coordinate responsible for a failure
Accepted scaffold edit
Edit magnitude under ⊕
Promotion scope: which loss is descended
on (USER, PROJECT, STACK, TENANT,
CORE)
Held-out patch eval before merge
Complexity penalty Ω ( z ) , bloat control,
cross-patch regression checks
Scaffold improves on this patch and
degrades elsewhere
Reverting a harmful accepted update

Loss

Gradient estimate

SGD step
Learning rate
Update radius

Validation loss
Regularization

Overfitting

Rollback

Credit assignment here is attribution, not backpropagation: the descent argument needs
only that a failure is traced to the responsible coordinate with non-zero probability, not
that the trace is computed by chain rule.

Appendix B. Survey detail

A representative spot-check of the corpus appears below, covering all six substrates, both
research and production, and all four clusters. PP is the integrated patch-local-adaptation
score (0–5); Evidence abbreviates the strongest source tier. The full corpus, the contrast
set, the gate-family mapping, and the per-system SkillOpt walkthrough are deferred to an
extended version.

System

Year

Substrate

Loop 1

Loop 2

PP

NanoResearch
SkillOpt

2026
2026

Auto
Auto

Auto
Auto

5
5

T2
T2

SkillRL
AlphaEvolve
DGM
SkillForge

2026
2025
2025
2026

Integrated
Integrated
(S2-led)
Integrated
Integrated
Integrated
Integrated

Auto
Auto
Auto
Auto

Auto
Auto
Auto
Human

5
4
4
4

T2
T2/T3
T2
T2

12

Evidence

## Page 13

Under review as a conference paper at COLM 2026

System

Year

Substrate

Loop 1

Loop 2

PP

Evidence

CASCADE
PACE

2025
2026

Auto
Auto

Human
Auto

4.5
4

T2
T2

CORAL

2026

Auto

Auto

4

T2

Meta-
Harness

2026

Auto

n/a

4

T2

Voyager
Claude
Code
CLAUDE.md
Cuadros
et
al. memory
MCP
Braintrust

2023
2025

Integrated
Integrated
(smallmodel)
Integrated
(S5+S3
async)
S5
Orchestration
Integrated
S1
Instructions
S3
Memory

Auto
n/a

Auto
n/a

4
4

T2
T3/T4

Auto

Auto

5

T2

n/a
n/a

n/a
Auto

5
5

T3
T3

Self-
Refine*
Mem0*

2023

S4 Tools
S6 Governance
(contrast)

n/a

n/a

1

T2

2024

(contrast)

n/a

n/a

2

T3/T4

2026

2024
2024

Asterisks mark contrast systems that fail the patch-local discriminator, included for definitional clarity.

The composite-two-loop systems sit in a 3 × 3 matrix indexed by gate type on each loop.
[Auto × Auto] holds the research full-auto systems with machine-evaluated criteria on both
loops. [Auto × Human] is the production-viable cell: Loop 1 fires at machine speed, Loop
2 requires human approval before propagation. [None × Auto] holds the governance-first
prompts-as-code systems with eval-gated CI but no failure-triggered candidate generation.
[Human × Auto] is empty by engineering logic (manual authorship offers little leverage for
automated promotion); [Human × Human] was empty in the surveyed window; [Human ×
None] collapses to “engineer edits a config” and is not patch-local in the sense of §3. The
empty cross-tenant governed-optimization corner of §6 is the gap, not any of these.

13
