{
  "slug": "frontier-and-localhost",
  "title": "Frontier and Localhost: How Production AI Learns Outside the Weights",
  "short": "FAL",
  "line": "errors",
  "line_name": "Error-accumulation",
  "part": 3,
  "status": "COLM 2026 workshop submission",
  "date": "2026-06-23",
  "authors": [
    "Mikhail L Arbuzov",
    "Sisong Bei",
    "Ziwei Dong",
    "Dmitri Kalaev",
    "Alexey Shvets"
  ],
  "abstract": "Production LLM systems increasingly adapt outside the model weights. After deployment failures, teams modify prompts, rules, memories, skills, tools, eval suites, routing graphs, and governance pipelines on a cadence that frontier weight updates cannot match. But this scaffold layer is still mostly maintained as patchwork: fixes are proposed by intuition, committed with weak credit assignment, accumulated without pruning, and promoted beyond the scope where they were validated. This paper formalises a disciplined alternative, which we call artifact-layer descent. The pipeline is concrete. Recurring residuals are assigned to scaffold coordinates. Candidate artifact deltas are tested against patch loss. Accepted deltas persist with rollback. Promotion across contexts is bounded by evidence radius: a local fix spreads only as far as the evidence supports. Surveying approximately 130 production and research systems from 2023 to 2026, we find current systems partially instantiate this loop but leave a central architecture gap, namely governed scaffold optimization across contexts. The contribution is not the claim that prompts are weights. It is to name the missing optimizer for an adaptation layer the field has already built.",
  "tldr": "Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations.",
  "pages": 13,
  "html": "https://telegrapher.ai/research/frontier-and-localhost/",
  "md": "https://telegrapher.ai/research/frontier-and-localhost.md",
  "reader": "https://telegrapher.ai/research/frontier-and-localhost/read/",
  "pdf": "https://telegrapher.ai/papers/frontier-and-localhost/frontier-and-localhost.pdf",
  "arxiv": null,
  "openreview": null,
  "post": "https://telegrapher.ai/blog/frontier-and-localhost/",
  "bibtex": "@misc{arbuzov2026frontier,\n  title         = {Frontier and Localhost: How Production AI Learns Outside the Weights},\n  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dong, Ziwei and Kalaev, Dmitri and Shvets, Alexey},\n  year          = {2026},\n  note          = {COLM 2026 workshop submission},\n  url           = {https://telegrapher.ai/research/frontier-and-localhost/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "Of ~130 LLM systems, none pairs automated scaffold fixes with governed cross-org promotion"
    },
    {
      "label": "TL;DR",
      "text": "Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations."
    },
    {
      "label": "Method",
      "text": "A survey of roughly 130 LLM systems published or productionised between January 2023 and May 2026, scored on a uniform 0–5 rubric across six scaffold substrates and tagged by evidence tier, read against a formal model of gated scaffold updates whose descent inequality and convergence proposition are stated in an appendix."
    },
    {
      "label": "Key result",
      "text": "LLM systems surveyed: ~130; Systems in the core corpus: ~90; Systems passing the composite two-loop audit: thirteen"
    },
    {
      "label": "Why it matters",
      "text": "For teams running agents in production, the practical change is small and specific: treat each rule, memory entry, skill or tool binding as an update with a measured effect, a rollback path and a scope, not as a note left for the next engineer."
    },
    {
      "label": "Limits",
      "text": "The evidence is a survey and a formal argument."
    },
    {
      "label": "Status",
      "text": "COLM 2026 workshop submission, June 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/frontier-and-localhost/read/, PDF /papers/frontier-and-localhost/frontier-and-localhost.pdf"
    }
  ],
  "note": {
    "slug": "frontier-and-localhost",
    "claim_title": "Of ~130 LLM systems, none pairs automated scaffold fixes with governed cross-org promotion",
    "meta_description": "LLM systems now adapt through prompts, memory, skills and tools. A survey of ~130 finds that loop run mostly as patchwork, its cross-org optimizer missing.",
    "tldr": "Production LLM systems are increasingly fixed by editing prompts, rules, memories, skills and tools, not weights. Across about 130 surveyed systems, that loop runs mostly as patchwork, and none governs its fixes across organisations.",
    "gist": "Production LLM systems increasingly adapt outside the model weights. After a failure, a team edits an instruction file, a memory, a skill, a tool binding, an eval or a routing graph, and these edits land far faster than frontier models are retrained. Most of it is patchwork. Fixes come from intuition, nobody tracks which edit helped, edits pile up unpruned, and a fix checked in one deployment drifts into others. The disciplined version, which the paper calls artifact-layer descent, clusters logged failures into recurring modes, traces a mode to the scaffold coordinate responsible and proposes one edit; a gate keeps it only if patch loss falls, a later check can roll it back, and promotion stops at the edit's evidence radius. Surveying roughly 130 systems from January 2023 to May 2026, scored on a uniform 0–5 rubric across six substrates, the paper finds the pieces of this loop scattered across systems but none that pairs an automated local gate with versioned, access-controlled, reversible promotion across organisations. The evidence is a survey and a formalisation.",
    "method": "A survey of roughly 130 LLM systems published or productionised between January 2023 and May 2026, scored on a uniform 0–5 rubric across six scaffold substrates and tagged by evidence tier, read against a formal model of gated scaffold updates whose descent inequality and convergence proposition are stated in an appendix.",
    "summary_html": [
      "The paper treats the layer around a frozen model as the thing being fit. It splits that scaffold into six <em>substrates</em> (instructions, skills, memory, tools, orchestration, governance), each an independently editable coordinate, and defines <em>patch loss</em> as the deployment's failure rate on its own recurring tasks. The disciplined loop, <em>artifact-layer descent</em>, starts from the accumulated failure record rather than from single incidents. It clusters failures into recurring modes and takes them in order of frequency times severity. For each mode it traces the coordinate responsible, such as a stale memory entry, a missing tool argument or a misrouted agent edge, and synthesises one candidate edit, which a gate keeps only if measured loss falls by a margin. A second loop decides how far an accepted edit travels. There the gate tightens as the <em>evidence radius</em> widens, because an edit that lowers a narrow average can raise a broad one. Against this model, roughly 130 systems from January 2023 to May 2026 are scored 0–5 per substrate; production claims are not pooled with peer-reviewed evidence.",
      "About 90 systems clear the bar for patch-local adaptation, and roughly 38 more that were described as self-improving, scaffolded or agentic but fall short of it are kept as a contrast set. Every mature system has a recognisable version of each substrate. None matures all six. Research systems learn fast and lack enterprise governance; production systems govern well and learn slowly; open standards such as AGENTS.md and MCP make artifacts portable without making them adaptive. A stricter two-loop audit is passed by thirteen research systems, and when the surveyed two-loop systems are indexed by the kind of gate on each loop, they fill three cells of a 3 × 3 matrix. The missing architecture is defined by four properties: an automated gate on local edits, promotion between organisations, versioned lineage with role-based access, and rollback of a bad promotion without a redeploy. The closest systems split these between them. No surveyed system holds all four, and none holds the cross-organisational three under any kind of gate."
    ],
    "key_numbers": [
      {
        "label": "LLM systems surveyed",
        "value": "~130",
        "context": "production and research systems published or productionised between January 2023 and May 2026"
      },
      {
        "label": "Systems in the core corpus",
        "value": "~90",
        "context": "score ≥ 3 on at least one of six substrates; roughly 38 more fail that bar and are kept as a contrast set"
      },
      {
        "label": "Systems passing the composite two-loop audit",
        "value": "thirteen",
        "context": "research systems whose Loop 1 edits a persisted artifact from session signal and whose Loop 2 promotes it under an explicit gate; the count moves under looser criteria"
      },
      {
        "label": "SkillOpt's average margin over the strongest baseline",
        "value": "+5.4 points",
        "context": "across 52 (model, benchmark, harness) cells, as reported by Yang et al. (2026) for a single skill document optimised as trainable external state"
      },
      {
        "label": "Gain from LLM-authored skills over no skills",
        "value": "+0.0 pp",
        "context": "against +16.2 pp for human-curated skills, as reported in work the paper cites; X. Zhang et al. (2026) attribute the gap to lifecycle management",
        "bad": true
      }
    ],
    "editorial_html": [
      "For teams running agents in production, the practical change is small and specific: treat each rule, memory entry, skill or tool binding as an update with a measured effect, a rollback path and a scope, not as a note left for the next engineer. Much of the machinery already exists under other names. An eval that blocks a merge is a validation check; a PR review against a shared rule repository is a promotion gate. What tends to be missing sits between the two, in tracing a failure back to the coordinate that caused it and in pruning edits that no longer earn their place.",
      "The framing also says where patchwork breaks. Instruction files bloat. Persistent memory can be poisoned by a single malicious entry that is read back long after the attack, tool outputs carry prompt injection across vendors, and the eval suite that gates edits drifts along with the scaffold it is supposed to judge. The paper treats each of these as a missing piece of the optimizer, not as hygiene for later.",
      "This is Part 3 of the error trilogy. <em>Beyond Exponential Decay</em> argues that long-context reliability hinges on a handful of decision points. <em>The Architecture of Errors</em> argues that inside an operationally bounded patch, failures fall into a small recurring catalogue, so reliability becomes a matter of discovering that catalogue and covering it. Frontier and Localhost follows the covering into production, where it happens in the scaffold. A residual made of a few repeating modes, each attachable to a scaffold coordinate, is what makes a local optimizer tractable at all; it also tells the optimizer where to spend."
    ],
    "limitations": "The evidence is a survey and a formal argument. Systems were found by searching public discourse and graded by evidence tier, from peer-reviewed papers down to industry-blog documentation. The two-loop count moves under two natural relaxations of the audit criterion, though the qualitative findings hold under both. The full corpus, the contrast set and the gate-family mapping are deferred to an extended version, so readers get a representative spot-check rather than every scored system. The descent reading is mechanism where its five conditions hold and analogy elsewhere. Where it applies, the convergence proposition promises a gate-stable scaffold rather than a global optimum, and in a deployment whose task mix shifts it tracks a moving optimum instead of settling. Whether the missing cross-organisational architecture should be built at all depends on the deployment; in safety-critical domains a reviewer's judgement in the gate may be a design feature rather than a defect. And the gates are themselves part of the moving scaffold, so keeping them held-out and recalibrated remains an open problem.",
    "concept_terms": [
      "artifact-layer descent",
      "evidence radius",
      "patch loss",
      "scaffold optimization",
      "two-loop architecture",
      "self-improving agents"
    ],
    "references": [
      "Tiwari et al. (2026). Learning, Fast and Slow: Towards LLMs That Adapt Continually.",
      "Kirk et al. (2024). Understanding the Effects of RLHF on LLM Generalisation and Diversity.",
      "Yang et al. (2026). SkillOpt: Executive Strategy for Self-Evolving Agent Skills.",
      "Agrawal et al. (2025). GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.",
      "Singh et al. (2025). The Leaderboard Illusion."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/frontier-and-localhost.txt",
    "verification": {
      "claims": [
        {
          "claim": "~130 LLM systems surveyed (claim_title, meta_description, tldr, gist, method, summary, key_numbers)",
          "verdict": "CONFIRMED",
          "source_quote": "The survey covers roughly 130 LLM systems published or productionised between January 2023 and May 2026"
        },
        {
          "claim": "Survey window January 2023 to May 2026",
          "verdict": "CONFIRMED",
          "source_quote": "published or productionised between January 2023 and May 2026"
        },
        {
          "claim": "Systems were found by searching public discourse for systems described as self-improving, scaffolded or agentic",
          "verdict": "CONFIRMED",
          "source_quote": "assembled by searching public discourse for systems described as"
        },
        {
          "claim": "About 90 systems form the core corpus: score ≥ 3 on at least one substrate (key_numbers, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "About 90 satisfy this patch-local discriminator (score ≥ 3 on at least one substrate) and form the core corpus"
        },
        {
          "claim": "Roughly 38 more carry the label but fail the bar and are kept as a contrast set",
          "verdict": "CONFIRMED",
          "source_quote": "roughly 38 carry the name but fail it and are kept as a contrast corpus"
        },
        {
          "claim": "Systems are tagged by evidence tier, from peer-reviewed papers down to industry-blog documentation",
          "verdict": "CONFIRMED",
          "source_quote": "from peer-reviewed publication (T1) down to industry-blog documentation (T5)"
        },
        {
          "claim": "Production claims are not pooled with peer-reviewed evidence",
          "verdict": "CONFIRMED",
          "source_quote": "production claims are not pooled with peer-reviewed evidence when computing cluster-level statistics"
        },
        {
          "claim": "Thirteen research systems pass the stricter composite two-loop audit (key_numbers, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "is passed by thirteen research systems"
        },
        {
          "claim": "Audit criterion: Loop 1 edits a persisted artifact from session signal; Loop 2 promotes under an explicit gate",
          "verdict": "CONFIRMED",
          "source_quote": "Loop 1 modifies a persisted artifact from session signal and Loop 2 has an explicit gated cross-context promotion mechanism"
        },
        {
          "claim": "The two-loop count moves under two relaxations; qualitative findings hold under both (key_numbers, limitations)",
          "verdict": "CONFIRMED",
          "source_quote": "the count moves under two natural relaxations of the criterion, but the qualitative findings below survive both"
        },
        {
          "claim": "Six substrates, each an independently editable coordinate",
          "verdict": "CONFIRMED",
          "source_quote": "They are the coordinates of θ s , each independently editable"
        },
        {
          "claim": "Six substrates are instructions, skills, memory, tools, orchestration, governance",
          "verdict": "CONFIRMED",
          "source_quote": "S1. Instructions"
        },
        {
          "claim": "Scored on a uniform 0–5 rubric (gist, method, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "The 0–5 rubric is uniform"
        },
        {
          "claim": "The scaffold, not the frozen frontier model, is the thing being fit",
          "verdict": "CONFIRMED",
          "source_quote": "with the frontier weights θ M fixed and the scaffold θ s — a structured object over the six substrate coordinates — the thing being fit"
        },
        {
          "claim": "Patch loss is the deployment's failure rate on its own recurring tasks",
          "verdict": "CONFIRMED",
          "source_quote": "deployment’s failure rate on its own recurring tasks"
        },
        {
          "claim": "Production LLM systems increasingly adapt outside the model weights",
          "verdict": "CONFIRMED",
          "source_quote": "Production LLM systems increasingly adapt outside the model weights."
        },
        {
          "claim": "Scaffold edits land on a cadence frontier weight updates cannot match (gist: 'far faster than frontier models are retrained')",
          "verdict": "CONFIRMED",
          "source_quote": "on a cadence that frontier weight updates cannot match"
        },
        {
          "claim": "The scaffold loop runs mostly as patchwork: fixes by intuition, weak credit assignment, no pruning, promoted beyond validated scope (gist, tldr, meta_description after correction)",
          "verdict": "CONFIRMED",
          "source_quote": "this scaffold layer is still mostly maintained as patchwork: fixes are proposed by intuition, committed with weak credit assignment, accumulated without pruning, and promoted beyond the scope where they were validated"
        },
        {
          "claim": "The disciplined loop works from the accumulated failure record, not single incidents",
          "verdict": "CONFIRMED",
          "source_quote": "It does not react to single failures; it works from their accumulated record."
        },
        {
          "claim": "Failures are clustered into recurring modes, taken in order of frequency times severity",
          "verdict": "CONFIRMED",
          "source_quote": "repeatability times severity, which is exactly its contribution to the deployment’s failure rate"
        },
        {
          "claim": "Coordinate examples: stale memory entry, missing tool argument, misrouted agent edge",
          "verdict": "CONFIRMED",
          "source_quote": "a stale memory entry, an under-specified instruction, a missing tool argument, a brittle retrieval rule, a misrouted agent edge"
        },
        {
          "claim": "One candidate edit per mode; a gate keeps it only if measured loss falls by a margin",
          "verdict": "CONFIRMED",
          "source_quote": "the gate accepts it when its noisy estimate of the directional loss change clears a margin"
        },
        {
          "claim": "A later check can roll back an accepted edit",
          "verdict": "CONFIRMED",
          "source_quote": "prunes or rolls back one that a later check shows made things worse"
        },
        {
          "claim": "A second loop decides how far an edit travels; the gate tightens as the evidence radius widens because a narrow-average gain can raise a broad average",
          "verdict": "CONFIRMED",
          "source_quote": "the gate gets stricter as the radius widens, because an edit that lowers a narrow average can raise a broad one"
        },
        {
          "claim": "Promotion stops at the edit's evidence radius",
          "verdict": "CONFIRMED",
          "source_quote": "promoting it to other deployments only as far as its evidence reaches — the edit’s evidence radius"
        },
        {
          "claim": "Every mature system has a recognisable version of each substrate; none matures all six",
          "verdict": "CONFIRMED",
          "source_quote": "every mature system has a recognisable instance of each, under different names — yet no system matures all six at once"
        },
        {
          "claim": "Research systems learn fast and lack governance; production systems govern well and learn slowly; open standards make artifacts portable, not adaptive",
          "verdict": "CONFIRMED",
          "source_quote": "research systems learn fast and lack enterprise governance, production systems govern well and learn slowly, and open standards solve artifact portability but not adaptation"
        },
        {
          "claim": "AGENTS.md and MCP are open standards",
          "verdict": "CONFIRMED",
          "source_quote": "Open standards (AGENTS.md, MCP, Claude Skills, Cursor Rules)"
        },
        {
          "claim": "The surveyed two-loop systems, indexed by gate type on each loop, fill three cells of a 3 × 3 matrix",
          "verdict": "CONFIRMED",
          "source_quote": "Indexed by gate type on each loop, the composite systems populate only three cells of a 3 × 3 matrix"
        },
        {
          "claim": "Four properties define the missing architecture: automated local gate, cross-org promotion, versioned lineage with role-based access, rollback without redeploy",
          "verdict": "CONFIRMED",
          "source_quote": "(b) multi-tenant cross-org promotion, Loop 2 routing artifacts between organisations; (c) versioned lineage with RBAC; and (d) rollback of a bad promotion without redeploy"
        },
        {
          "claim": "The closest systems split the four properties between them",
          "verdict": "CONFIRMED",
          "source_quote": "The closest approaches partition the requirements."
        },
        {
          "claim": "No surveyed system holds all four; the cross-organisational three are unfilled under any gate (claim_title, tldr, gist, summary)",
          "verdict": "CONFIRMED",
          "source_quote": "No surveyed system combines all four, and the cross-organisational triad (b, c, d) is itself unfilled under any gate type"
        },
        {
          "claim": "SkillOpt: average margin +5.4 points over the strongest baseline (key_numbers)",
          "verdict": "CONFIRMED",
          "source_quote": "beating the strongest per-cell baseline by +5.4 points on average"
        },
        {
          "claim": "SkillOpt evaluated on 52 (model, benchmark, harness) cells",
          "verdict": "CONFIRMED",
          "source_quote": "best or tied-best on all 52 evaluated (model, benchmark, harness) cells"
        },
        {
          "claim": "SkillOpt (Yang et al. 2026) optimises a single skill document as trainable external state",
          "verdict": "CONFIRMED",
          "source_quote": "treats a single markdown skill document as the trainable external state of a frozen agent"
        },
        {
          "claim": "LLM-authored skills +0.0 pp over no skills vs +16.2 pp for human-curated skills, as reported in cited work (key_numbers, bad)",
          "verdict": "CONFIRMED",
          "source_quote": "LLM-authored skills have been reported at + 0.0 pp over a no-skill baseline against + 16.2 pp for human-curated ones"
        },
        {
          "claim": "X. Zhang et al. (2026) attribute the gap to lifecycle management",
          "verdict": "CONFIRMED",
          "source_quote": "a gap (X. Zhang et al. 2026) attributes to lifecycle management"
        },
        {
          "claim": "Editorial: an eval that blocks a merge is a validation check; a PR review against a shared rule repository is a promotion gate",
          "verdict": "CONFIRMED",
          "source_quote": "A PR review against a shared rule repository is a promotion gate on a federated update. An eval suite that blocks merge is a validation-loss check before promotion."
        },
        {
          "claim": "Editorial: much of the machinery already exists under other names",
          "verdict": "CONFIRMED",
          "source_quote": "None of these mechanisms is new"
        },
        {
          "claim": "Editorial: what tends to be missing is failure-to-coordinate tracing and pruning",
          "verdict": "CONFIRMED",
          "source_quote": "Eval-gated CI, no failure-to-coordinate tracing"
        },
        {
          "claim": "Editorial: instruction files bloat",
          "verdict": "CONFIRMED",
          "source_quote": "instruction-following degrades and latency grows sharply as instruction count scales"
        },
        {
          "claim": "Editorial: one poisoned memory entry can be read back long after the attack",
          "verdict": "CONFIRMED",
          "source_quote": "one malicious entry can be retrieved and acted on long after the attack"
        },
        {
          "claim": "Editorial: tool outputs carry prompt injection across vendors",
          "verdict": "CONFIRMED",
          "source_quote": "Prompt injection via tools and MCP makes a tool’s output a cross-vendor attack surface"
        },
        {
          "claim": "Editorial: the eval suite drifts along with the scaffold it gates",
          "verdict": "CONFIRMED",
          "source_quote": "because the suite adapts on the same cadence as the scaffold it gates"
        },
        {
          "claim": "Editorial: each failure mode is a missing piece of the optimizer, not later hygiene",
          "verdict": "CONFIRMED",
          "source_quote": "Each one is also a missing piece of the optimizer."
        },
        {
          "claim": "Editorial: a residual of a few repeating modes, each attachable to a coordinate, makes a local optimizer tractable and tells it where to spend",
          "verdict": "CONFIRMED",
          "source_quote": "A patch’s residual is therefore not diffuse noise but a small set of repeating modes, each attachable to a scaffold coordinate — which is what makes a disciplined local optimizer possible at all, and what tells that optimizer where to spend its budget"
        },
        {
          "claim": "Editorial: the paper builds on earlier error-accumulation papers by the same authors",
          "verdict": "CONFIRMED",
          "source_quote": "Prior work on error accumulation (Arbuzov et al. 2025, 2026)"
        },
        {
          "claim": "Editorial: this is Part 3 of the error trilogy (site metadata; the paper text itself does not mention the trilogy)",
          "verdict": "CONFIRMED",
          "source_quote": "[content/papers.yaml] slug: frontier-and-localhost ... line: errors part: 3"
        },
        {
          "claim": "Editorial: Beyond Exponential Decay argues long-context reliability hinges on a handful of decision points (checked against that paper's abstract)",
          "verdict": "CONFIRMED",
          "source_quote": "[content/abstracts.yaml] long-context reliability hinges on a handful of decision points"
        },
        {
          "claim": "Editorial: The Architecture of Errors argues failures inside an operationally bounded patch fall into a small recurring catalogue, making reliability a catalogue-discovery and coverage problem (checked against that paper's abstract)",
          "verdict": "CONFIRMED",
          "source_quote": "[content/abstracts.yaml] Within such patches, empirical evidence suggests failures are sparse, repetitive, and concentrated in a small recurring catalogue, so reliability becomes a local catalogue-discovery and intervention-coverage problem"
        },
        {
          "claim": "Limitations/method: the evidence is a survey and a formal argument; descent inequality and convergence proposition stated in an appendix",
          "verdict": "CONFIRMED",
          "source_quote": "Appendix A states the assumptions, a one-step descent inequality, and a convergence proposition to a gate-stable scaffold"
        },
        {
          "claim": "Limitations: descent reading is mechanism where the five conditions hold, analogy elsewhere",
          "verdict": "CONFIRMED",
          "source_quote": "The descent reading is mechanism when the five conditions above hold and analogy otherwise"
        },
        {
          "claim": "Limitations: full corpus, contrast set and gate-family mapping deferred; only a representative spot-check",
          "verdict": "CONFIRMED",
          "source_quote": "The full corpus, the contrast set, the gate-family mapping, and the per-system SkillOpt walkthrough are deferred to an extended version"
        },
        {
          "claim": "Limitations: convergence is to a gate-stable scaffold, not a global optimum",
          "verdict": "CONFIRMED",
          "source_quote": "not convergence to a global optimum"
        },
        {
          "claim": "Limitations: under a shifting task mix the loop tracks a moving optimum",
          "verdict": "CONFIRMED",
          "source_quote": "the same mechanism tracks a moving optimum rather than converging once"
        },
        {
          "claim": "Limitations: whether to build the missing architecture depends on the deployment; a reviewer's judgement in the gate may be a design feature in safety-critical domains (corrected from 'left open')",
          "verdict": "CONFIRMED",
          "source_quote": "Whether it should be filled is deployment-dependent — in safety-critical domains a reviewer-judgment gate may be a design feature rather than a defect"
        },
        {
          "claim": "Limitations: the gates are part of the moving scaffold and need held-out recalibration",
          "verdict": "CONFIRMED",
          "source_quote": "The instruments that validate these edits are themselves part of the moving, clustered scaffold"
        },
        {
          "claim": "Reference present: Tiwari et al. (2026) Learning, Fast and Slow",
          "verdict": "CONFIRMED",
          "source_quote": "Learning, Fast and Slow: Towards LLMs That Adapt Continually"
        },
        {
          "claim": "Reference present: Kirk et al. (2024) RLHF generalisation and diversity",
          "verdict": "CONFIRMED",
          "source_quote": "Understanding the Effects of RLHF on LLM Generalisation and Diversity"
        },
        {
          "claim": "Reference present: Yang et al. (2026) SkillOpt",
          "verdict": "CONFIRMED",
          "source_quote": "SkillOpt: Executive Strategy for Self-Evolving Agent Skills"
        },
        {
          "claim": "Reference present: Agrawal et al. (2025) GEPA",
          "verdict": "CONFIRMED",
          "source_quote": "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning"
        },
        {
          "claim": "Reference present: Singh et al. (2025) The Leaderboard Illusion",
          "verdict": "CONFIRMED",
          "source_quote": "The Leaderboard Illusion"
        },
        {
          "claim": "CORRECTED (tldr, meta_description): 'that loop runs as patchwork' across the ~130 surveyed systems dropped the paper's 'mostly'; the paper says the layer is 'mostly maintained as patchwork' and the survey finds systems 'partially instantiate this loop' (thirteen research systems close both loops). Now 'mostly as patchwork'.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (gist, method): 'each scored 0–5 on six (scaffold) substrates' implied six scores per system; the substrate table counts systems per substrate (n from 12 to 34, summing past the corpus size). Reworded to 'scored on a uniform 0–5 rubric across six substrates'.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (summary_html[1]): the ~38 contrast systems 'carry the self-improving label'; the paper searched for systems described as 'self-improving', 'scaffolded' or 'agentic', and the contrast set 'carry the name'. Reworded to name all three labels.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        },
        {
          "claim": "CORRECTED (limitations): 'whether the missing architecture should be built at all is left open'; the paper says it is 'deployment-dependent'. Reworded.",
          "verdict": "UNVERIFIED",
          "source_quote": ""
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}