{
  "slug": "ripplekb",
  "title": "RippleKB: Finding What an Edit Changes Across Linked Documents",
  "short": "RKB",
  "line": "telegraph",
  "line_name": "Telegraph English",
  "part": null,
  "status": "ICLR 2027 submission",
  "date": "2026-08-25",
  "authors": [
    "Sisong Bei",
    "Mikhail L Arbuzov",
    "Ziwei Dong",
    "Alexey Shvets",
    "Dmitri Kalaev"
  ],
  "abstract": "Editing one fact can change a derived quantity, comparison, or other statement across linked documents. RippleKB evaluates whether systems can recover a constructed set of affected source units completely. Each item edits one source fact in a small set of linked documents. The proposed answer key is a hidden dependency closure, with unaffected units also present. A system returns a ranked list of source units, word for word, under a review rule tied to the constructed affected-set size. We regenerate every item from a public multi-hop corpus and test identifier-order, document-position, lexical, and supporting-fact rankings, alongside role and template concentration checks. We release the audit protocol under which the proposed closure is to be reviewed unit by unit. We report how far retrieval-style and language-model baselines get toward recovering the constructed closure completely on the machine-validated candidate set. Complete recovery tests whether a ranking includes the entire constructed target, while exact-copy measures separately assess preservation of source text.",
  "tldr": "Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.",
  "pages": 21,
  "html": "https://telegrapher.ai/research/ripplekb/",
  "md": "https://telegrapher.ai/research/ripplekb.md",
  "reader": "https://telegrapher.ai/research/ripplekb/read/",
  "pdf": "https://telegrapher.ai/papers/ripplekb/ripplekb.pdf",
  "arxiv": null,
  "openreview": "https://openreview.net/forum?id=zRZ3MrSLIj",
  "post": "https://telegrapher.ai/blog/ripplekb/",
  "bibtex": "@misc{bei2026ripplekb,\n  title         = {RippleKB: Finding What an Edit Changes Across Linked Documents},\n  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Shvets, Alexey and Kalaev, Dmitri},\n  year          = {2026},\n  note          = {ICLR 2027 submission},\n  url           = {https://telegrapher.ai/research/ripplekb/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "A ranking can recover most of what an edit changed and still miss the edited sentence"
    },
    {
      "label": "TL;DR",
      "text": "Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself."
    },
    {
      "label": "Method",
      "text": "150 machine-validated items regenerated from MuSiQue, each a numeric edit over three or four linked documents with a hidden dependency closure of four to seven units, were ranked by deterministic baselines (random, render and identifier order, token overlap, BM25), BM25 with Titan Text Embeddings V2 reranking, Qwen3 32B, Llama 3.3 70B and Claude Sonnet 4.5, plus four hosted models added after the initial results were seen (Claude Sonnet 5, Claude Opus 5, Claude Fable 5.1, Grok 4.6), and scored by complete-closure@2G, Recall@2G, Recall@G and exact-copy fidelity under a no-repair parser."
    },
    {
      "label": "Key result",
      "text": "Complete recovery, Claude Opus 5: 0%; Complete recovery, BM25 with embedding reranking: 13%; Complete recovery, Llama 3.3 70B: 65%"
    },
    {
      "label": "Why it matters",
      "text": "This is for people who maintain text that refers to other text: totals computed from figures held elsewhere, summaries that restate numbers from source pages, corpora that downstream answers lean on."
    },
    {
      "label": "Limits",
      "text": "The answer key is the construction's own."
    },
    {
      "label": "Status",
      "text": "ICLR 2027 submission, August 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/ripplekb/read/, PDF /papers/ripplekb/ripplekb.pdf"
    }
  ],
  "note": {
    "slug": "ripplekb",
    "claim_title": "A ranking can recover most of what an edit changed and still miss the edited sentence",
    "meta_description": "RippleKB scores whether a ranking recovers every statement one edit changes across linked documents. Recall and complete recovery rank systems differently.",
    "tldr": "Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.",
    "gist": "Correct one figure in a document and a total in another document can change, along with a comparison built on that total; find the edited sentence alone and the rest goes unreviewed. RippleKB turns this into a ranking task. Each of 150 items, regenerated from MuSiQue and machine-validated, pairs a visible numeric edit with three or four linked documents, where a hidden dependency closure marks four to seven affected units among 14–26 candidates. A system ranks the units, copying each word for word. The scorer reads a cutoff of twice the affected-set size, which the system does not see, and takes two scores from it: Recall@2G, the mean fraction found, and complete-closure@2G, the share of items found in full. They diverge. Embedding reranking moved BM25's recall point estimate from 0.66 to 0.68 while complete recovery fell from 20% to 13%. Claude Opus 5 reached 0.79 recall and completed no item, because the edited sentence stayed outside its cutoff. The answer key is construction-defined; a human review protocol is released, with no review results reported.",
    "method": "150 machine-validated items regenerated from MuSiQue, each a numeric edit over three or four linked documents with a hidden dependency closure of four to seven units, were ranked by deterministic baselines (random, render and identifier order, token overlap, BM25), BM25 with Titan Text Embeddings V2 reranking, Qwen3 32B, Llama 3.3 70B and Claude Sonnet 4.5, plus four hosted models added after the initial results were seen (Claude Sonnet 5, Claude Opus 5, Claude Fable 5.1, Grok 4.6), and scored by complete-closure@2G, Recall@2G, Recall@G and exact-copy fidelity under a no-repair parser.",
    "summary_html": [
      "In the paper's conceptual example, one sentence says the North store holds <em>n</em> crates, a second gives combined stock as <em>n + m</em>, and a third says combined stock is below <em>q</em>. Raise North's count far enough for the total to reach <em>q</em> and all three change. Two neighbours do not: South's count feeds the total but stays put, and a note that North opens at dawn shares the subject and stays true. The changed sentences form the item's <em>impact set</em>, which RippleKB proposes through a hidden <em>dependency closure</em>. Its 150 machine-validated items, regenerated from MuSiQue, each edit one numeric value across three or four linked documents of 14–26 candidate units, four to seven of them affected. A system sees the edit and the documents, then ranks units, giving each one's identifier and its text copied byte for byte. The scorer reads the top <em>K</em> = min(|U|, 2|G|) positions, a budget set by the hidden affected-set size, and reports two numbers from that one prefix: <b>Recall@2G</b>, the mean fraction found, and <b>complete-closure@2G</b>, the share of items found whole. Four shortcut rankings were tested against the key (identifier order, document position, lexical overlap with the edit, the original question's supporting facts), and lexical bridging came closest, completing at most 20% of the pooled items.",
      "Recall and completion separate the systems, sometimes in opposite order. Under the doubled budget, random ordering already earns Recall@2G 0.54 while completing 1% of items. Reranking BM25 with embeddings raised the recall point estimate slightly and cut complete recovery from 20% to 13%. The stronger language-model rankings lift both: Llama 3.3 70B completed 65% of items at recall 0.91, and Claude Sonnet 5, one of four models added after the initial results were seen, 82% at 0.95. The gap is widest for Claude Opus 5, also a later addition, with recall 0.79 and no completed item; within its cutoff, it recovered the directly edited sentence on none of the items. Delivery is a separate loss. Claude Sonnet 4.5 wrapped every response in a code fence and scored zero under the no-repair parser, while the same saved responses, stripped of the fence, give 33% completion and recall 0.85. Copying is a third. Conditional exactness counts recovered units alone and runs high, from 0.942 for Llama upward among systems with parsed rankings. Opus's 0.999 on that measure sits beside numeric-span preservation of 0.28, because the span score also counts the units it left out."
    ],
    "key_numbers": [
      {
        "label": "Complete recovery, Claude Opus 5",
        "value": "0%",
        "context": "with Recall@2G 0.79; its recovery of the directly edited sentence within the review cutoff was zero; added after the initial results were seen",
        "bad": true
      },
      {
        "label": "Complete recovery, BM25 with embedding reranking",
        "value": "13%",
        "context": "against 20% for BM25 alone, while the Recall@2G point estimate rose from 0.66 to 0.68",
        "bad": true
      },
      {
        "label": "Complete recovery, Llama 3.3 70B",
        "value": "65%",
        "context": "Recall@2G 0.91; the later-added Claude Sonnet 5 reached 82% at 0.95"
      },
      {
        "label": "Recall@2G, random ordering",
        "value": "0.54",
        "context": "with complete-closure@2G of 1%; the doubled review budget alone hands out substantial partial credit"
      },
      {
        "label": "Complete recovery, Claude Sonnet 4.5, primary parser",
        "value": "0%",
        "context": "every response arrived in a code fence; stripping the fence from the same responses gives 33% and Recall@2G 0.85",
        "bad": true
      }
    ],
    "editorial_html": [
      "This is for people who maintain text that refers to other text: totals computed from figures held elsewhere, summaries that restate numbers from source pages, corpora that downstream answers lean on. After a correction, the operational question is whether anything affected is still unreviewed. Average recall does not answer it. A reranker that edges recall up can finish fewer review sets, and a model can post a respectable recall while leaving the very sentence that was edited outside its review cutoff. Report complete recovery beside recall, at a stated review budget, and both failures show.",
      "The paper also keeps apart failures that a single score would merge. An unparsable response counts as an empty ranking, so delivery is part of what is measured; harsh, but a pipeline that cannot read a ranking has no ranking. The fence-stripping reparse of the same responses is reported beside it as a diagnosis, not a replacement. Selection and copying are scored apart too, since a correct identifier can carry altered text and perfect copies can still leave members of the set missing.",
      "RippleKB sits with the Telegraph English papers through its unit of account. <em>Telegraph English</em> rewrites text into lines that are each an independently addressable fact. A store of addressable facts invites edits one fact at a time, and RippleKB scores the question that follows: which other units did the edit change? Its units are sentences in constructed prose documents rather than TE lines, so the link is a shared question, not a shared result. The habit of measuring separately also runs through <em>What Survives Learned Symbolic Compression?</em>, which treats source fidelity, internal validity and what a consumer model recovers as different quantities."
    ],
    "limitations": "The answer key is the construction's own. Scores measure recovery of a generated dependency closure on machine-validated candidates, and the released protocol for human review, which checks each proposed dependency against the text and looks for affected statements the closure missed, has no reported results. Edits are numeric, and each item is a small, finite collection of constructed English documents derived from Wikipedia material. The review cutoff uses the hidden affected-set size; a deployed reviewer would have to choose a cutoff without the answer key. Four hosted models were added after the initial results were seen and ran with provider-default sampling, and no provider run was repeated to estimate generation variability, so the intervals reflect source sampling alone. The evaluation pools all 150 candidates where the original plan named the report partition alone, and the construction-control report records access to selection and report labels before baseline outputs were available. Affected downstream sentences open with their named entity, a possible cue the shortcut rankings do not test. Prior model exposure to the source passages is unassessed, and independent replay needs item-level responses and executable scoring code beyond the aggregate materials. Grok 4.6 returned five parsable responses, so its row says little about how it ranks.",
    "concept_terms": [
      "change-impact analysis",
      "edit propagation",
      "impact set",
      "dependency closure",
      "complete recovery",
      "verbatim retrieval"
    ],
    "references": [
      "Trivedi et al. (2022). MuSiQue: Multihop Questions via Single-hop Question Composition.",
      "Goknil et al. (2016). A Rule-Based Change Impact Analysis Approach in Software Architecture for Requirements Changes.",
      "Cohen et al. (2024). Evaluating the Ripple Effects of Knowledge Editing in Language Models.",
      "Zhong et al. (2023). MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions.",
      "Logan IV et al. (2022). FRUIT: Faithfully Reflecting Updated Information in Text."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/ripplekb.txt",
    "verification": {
      "claims": [
        {
          "claim": "Dataset: 150 machine-validated candidate items",
          "verdict": "CONFIRMED",
          "source_quote": "The collection contains 150 machine-validated candidate items."
        },
        {
          "claim": "Items are regenerated from MuSiQue, whose questions compose facts across Wikipedia passages",
          "verdict": "CONFIRMED",
          "source_quote": "Each item is regenerated from MuSiQue, whose questions compose facts across Wikipedia passages"
        },
        {
          "claim": "Each edit changes one numeric value in a source sentence",
          "verdict": "CONFIRMED",
          "source_quote": "The edit changes a numeric value in a source-grounded sentence."
        },
        {
          "claim": "Items contain three or four documents, 14–26 candidate units, four to seven affected units",
          "verdict": "CONFIRMED",
          "source_quote": "Candidate items contain three or four documents and 14–26 units, with proposed target sets of four to seven units"
        },
        {
          "claim": "The answer key is a hidden dependency closure; graph and impact-set size are hidden from the system",
          "verdict": "CONFIRMED",
          "source_quote": "The dependency graph, impact-set size, source-record identifier, and split assignment remain hidden during prediction."
        },
        {
          "claim": "Systems rank units, returning identifier plus text copied byte for byte (word for word)",
          "verdict": "CONFIRMED",
          "source_quote": "Copy each unit's text byte-for-byte from the input."
        },
        {
          "claim": "Cutoff K = min(|U|, 2|G|), set by the hidden affected-set size; equals twice the affected-set size on every item (derived: |U| >= 14 >= 2|G| since |G| <= 7)",
          "verdict": "CONFIRMED",
          "source_quote": "K = min(|U |, 2|G|). The scorer knows |G|; the system ranks candidates without seeing that size or its resulting cutoff."
        },
        {
          "claim": "Recall@2G = mean fraction found; complete-closure@2G = share of items found whole; both from the same prefix",
          "verdict": "CONFIRMED",
          "source_quote": "Recall@2G is the mean fraction of affected identifiers recovered within K positions; complete-closure@2G is the proportion of items whose entire impact set is recovered there."
        },
        {
          "claim": "Conceptual example: North n crates, combined n + m, below q; South and 'opens at dawn' unchanged",
          "verdict": "CONFIRMED",
          "source_quote": "Unchanged: D South stores m crates. E North opens at dawn."
        },
        {
          "claim": "Finding the edited sentence alone leaves consequences unreviewed",
          "verdict": "CONFIRMED",
          "source_quote": "Finding the edited sentence alone leaves those consequences unreviewed."
        },
        {
          "claim": "Shortcut rankings tested: identifier order, document position, lexical overlap, supporting facts",
          "verdict": "CONFIRMED",
          "source_quote": "test identifier-order, document-position, lexical, and supporting-fact rankings"
        },
        {
          "claim": "Lexical bridging is the strongest shortcut, completing at most 20% of pooled items",
          "verdict": "CONFIRMED",
          "source_quote": "Lexical bridging is the strongest tested shortcut, completing at most 20% of pooled candidates"
        },
        {
          "claim": "Random ordering: Recall@2G 0.54, complete-closure@2G 1%",
          "verdict": "CONFIRMED",
          "source_quote": "The retained random baseline reaches Recall@2G 0.54 alongside complete-closure@2G 1%"
        },
        {
          "claim": "The doubled budget gives random ordering substantial partial credit",
          "verdict": "CONFIRMED",
          "source_quote": "The larger budget gives random selection substantial partial credit on the current candidate collection."
        },
        {
          "claim": "BM25 with embedding reranking: recall point estimate 0.66 -> 0.68, complete recovery 20% -> 13%",
          "verdict": "CONFIRMED",
          "source_quote": "Embedding reranking raises BM25’s partial-recall point estimate from 0.66 to 0.68 while lowering complete recovery from 20% to 13%."
        },
        {
          "claim": "Reranker is Titan Text Embeddings V2 over BM25 candidates",
          "verdict": "CONFIRMED",
          "source_quote": "The retrieval baseline reranks BM25 candidates using Titan Text Embeddings V2."
        },
        {
          "claim": "Initial model baselines: Qwen3 32B, Llama 3.3 70B, Claude Sonnet 4.5",
          "verdict": "CONFIRMED",
          "source_quote": "Qwen3 32B, Llama 3.3 70B, and Claude Sonnet 4.5 receive the edit and documents"
        },
        {
          "claim": "Deterministic baselines: random, render order, identifier order, token overlap, BM25",
          "verdict": "CONFIRMED",
          "source_quote": "Deterministic methods use random, rendered, or identifier order, token overlap, or BM25"
        },
        {
          "claim": "Four hosted models (Claude Sonnet 5, Claude Opus 5, Claude Fable 5.1, Grok 4.6) added after the first results were seen, with provider-default sampling",
          "verdict": "CONFIRMED",
          "source_quote": "Claude Sonnet 5, Claude Opus 5, Claude Fable 5.1, and Grok 4.6 were added after the first results were seen."
        },
        {
          "claim": "Stronger language-model rankings lift both recall and completion",
          "verdict": "CONFIRMED",
          "source_quote": "Thus the stronger language-model rankings improve both the amount of affected material found and the frequency of finishing the target."
        },
        {
          "claim": "Llama 3.3 70B: complete-closure@2G 65%, Recall@2G 0.91",
          "verdict": "CONFIRMED",
          "source_quote": "Llama 3.3 70B reaches complete-closure@2G 65% with Recall@2G 0.91"
        },
        {
          "claim": "Claude Sonnet 5 (later-added): 82% completion at Recall@2G 0.95",
          "verdict": "CONFIRMED",
          "source_quote": "The later-added Sonnet 5 reaches 82% and 0.95, the highest retained completion estimate."
        },
        {
          "claim": "Claude Opus 5 (later-added): Recall@2G 0.79, complete-closure@2G 0%",
          "verdict": "CONFIRMED",
          "source_quote": "Opus reaches Recall@2G 0.79 with complete-closure@2G 0%"
        },
        {
          "claim": "Opus recovered the directly edited sentence on no item within its review cutoff, which prevents completion",
          "verdict": "CONFIRMED",
          "source_quote": "its directunit recovery within the review prefix is zero"
        },
        {
          "claim": "Recall-completion gap is widest for Opus (derived from Table 10: Opus 0.79 vs next-widest render order 0.57)",
          "verdict": "CONFIRMED",
          "source_quote": "Claude Opus 5 0% [0, 0] 0.79 [0.78, 0.81]"
        },
        {
          "claim": "Claude Sonnet 4.5 fenced every response and scored zero under the no-repair primary parser",
          "verdict": "CONFIRMED",
          "source_quote": "Sonnet 4.5 fences every response and therefore receives complete-closure@2G 0% with Recall@2G 0.00."
        },
        {
          "claim": "Fence-stripped reparse of the same Sonnet 4.5 responses: 33% completion, Recall@2G 0.85",
          "verdict": "CONFIRMED",
          "source_quote": "The separately reported fence-stripping parse of those same responses gives 33% and 0.85."
        },
        {
          "claim": "An unparsable response is scored as an empty ranking; the reparse is a diagnosis, not a replacement",
          "verdict": "CONFIRMED",
          "source_quote": "It diagnoses a transport-format effect without replacing the primary scores or generating new responses."
        },
        {
          "claim": "Conditional exactness counts recovered units only; lowest among systems with parsed rankings is Llama's 0.942 (Table 15)",
          "verdict": "CONFIRMED",
          "source_quote": "Among recovered affected units, Llama’s conditional verbatim exactness is 0.942"
        },
        {
          "claim": "Opus conditional exactness 0.999 vs numeric-span preservation 0.28; span score includes units not returned",
          "verdict": "CONFIRMED",
          "source_quote": "Opus’s near-perfect conditional copying coexists with numeric-span preservation of 0.28 (Table 3)."
        },
        {
          "claim": "A correct identifier can carry altered text; perfect copies can still leave set members missing",
          "verdict": "CONFIRMED",
          "source_quote": "a correctly selected identifier can carry altered text, while perfectly copied units can still leave other members of the impact set missing."
        },
        {
          "claim": "Telegraph English (other paper): output lines are each an independently addressable fact",
          "verdict": "CONFIRMED",
          "source_quote": "abstracts.yaml telegraph-english: each output line is an independently addressable fact"
        },
        {
          "claim": "What Survives Learned Symbolic Compression? (other paper): source fidelity, internal validity and consumer recovery are different quantities",
          "verdict": "CONFIRMED",
          "source_quote": "abstracts.yaml what-survives-learned-symbolic-compression: Source fidelity relative to the constructed reference, internal validity, and what a consumer model recovers are three different quantities"
        },
        {
          "claim": "Limitation: answer key is construction-defined; human-review protocol released with no reported results",
          "verdict": "CONFIRMED",
          "source_quote": "the released human-review protocol specifies how to assess it."
        },
        {
          "claim": "Limitation: numeric edits in finite collections of constructed English documents derived from Wikipedia material",
          "verdict": "CONFIRMED",
          "source_quote": "RippleKB currently studies numeric edits in finite collections of constructed English documents derived from Wikipedia material."
        },
        {
          "claim": "Limitation: deployed review would need a cutoff chosen without the answer key",
          "verdict": "CONFIRMED",
          "source_quote": "deploying the task in a review workflow would also require choosing a cutoff without access to the answer key."
        },
        {
          "claim": "Limitation: provider runs not repeated to estimate generation variability; intervals reflect source sampling (corrected from 'no provider run was repeated': Appendix F records provider-compatibility reruns)",
          "verdict": "CONFIRMED",
          "source_quote": "provider runs were not repeated to estimate sampling variability from generation."
        },
        {
          "claim": "Limitation: evaluation pools all 150 candidates whereas the original plan specified a report-only population",
          "verdict": "CONFIRMED",
          "source_quote": "The pooled evaluation uses all construction partitions, whereas the original plan specified a report-only population."
        },
        {
          "claim": "Limitation: construction-control report records access to selection and report labels before baseline outputs",
          "verdict": "CONFIRMED",
          "source_quote": "The construction-control report records access to selection and report labels before baseline outputs were available."
        },
        {
          "claim": "Limitation: entity-first openings of affected downstream units are a potential cue the rankings do not test",
          "verdict": "CONFIRMED",
          "source_quote": "the entity-first openings of affected downstream units remain a potential cue outside these rankings."
        },
        {
          "claim": "Limitation: prior model exposure unassessed; independent replay needs item-level responses and executable scoring code",
          "verdict": "CONFIRMED",
          "source_quote": "shared passages and prior model exposure remain unassessed."
        },
        {
          "claim": "Limitation: Grok 4.6 supplied five parsed responses",
          "verdict": "CONFIRMED",
          "source_quote": "The later-added Grok run supplies only five parsed responses under the 4,096-token output cap"
        },
        {
          "claim": "Editorial (corrected): Opus left the edited sentence outside its review cutoff (was 'leaving out'; the paper's zero concerns only the reviewed prefix)",
          "verdict": "CONFIRMED",
          "source_quote": "Opus’s zero direct-recovery score concerns only positions within the reviewed prefix."
        },
        {
          "claim": "References: Trivedi 2022, Goknil 2016, Cohen 2024, Zhong 2023, Logan IV 2022 are all in the paper's reference list",
          "verdict": "CONFIRMED",
          "source_quote": "MuSiQue: Multihop Questions via Single-hop Question Composition."
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}