A ranking can recover most of what an edit changed and still miss the edited sentence
RippleKB: Finding What an Edit Changes Across Linked Documents
Sisong Bei, Mikhail L Arbuzov, Ziwei Dong, Alexey Shvets, Dmitri Kalaev
ICLR 2027 submission, August 2026
Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.
What we did and found
In the paper's conceptual example, one sentence says the North store holds n crates, a second gives combined stock as n + m, and a third says combined stock is below q. Raise North's count far enough for the total to reach q and all three change. Two neighbours do not: South's count feeds the total but stays put, and a note that North opens at dawn shares the subject and stays true. The changed sentences form the item's impact set, which RippleKB proposes through a hidden dependency closure. Its 150 machine-validated items, regenerated from MuSiQue, each edit one numeric value across three or four linked documents of 14–26 candidate units, four to seven of them affected. A system sees the edit and the documents, then ranks units, giving each one's identifier and its text copied byte for byte. The scorer reads the top K = min(|U|, 2|G|) positions, a budget set by the hidden affected-set size, and reports two numbers from that one prefix: Recall@2G, the mean fraction found, and complete-closure@2G, the share of items found whole. Four shortcut rankings were tested against the key (identifier order, document position, lexical overlap with the edit, the original question's supporting facts), and lexical bridging came closest, completing at most 20% of the pooled items.
Recall and completion separate the systems, sometimes in opposite order. Under the doubled budget, random ordering already earns Recall@2G 0.54 while completing 1% of items. Reranking BM25 with embeddings raised the recall point estimate slightly and cut complete recovery from 20% to 13%. The stronger language-model rankings lift both: Llama 3.3 70B completed 65% of items at recall 0.91, and Claude Sonnet 5, one of four models added after the initial results were seen, 82% at 0.95. The gap is widest for Claude Opus 5, also a later addition, with recall 0.79 and no completed item; within its cutoff, it recovered the directly edited sentence on none of the items. Delivery is a separate loss. Claude Sonnet 4.5 wrapped every response in a code fence and scored zero under the no-repair parser, while the same saved responses, stripped of the fence, give 33% completion and recall 0.85. Copying is a third. Conditional exactness counts recovered units alone and runs high, from 0.942 for Llama upward among systems with parsed rankings. Opus's 0.999 on that measure sits beside numeric-span preservation of 0.28, because the span score also counts the units it left out.
Key numbers
| Complete recovery, Claude Opus 5with Recall@2G 0.79; its recovery of the directly edited sentence within the review cutoff was zero; added after the initial results were seen | 0% |
| Complete recovery, BM25 with embedding rerankingagainst 20% for BM25 alone, while the Recall@2G point estimate rose from 0.66 to 0.68 | 13% |
| Complete recovery, Llama 3.3 70BRecall@2G 0.91; the later-added Claude Sonnet 5 reached 82% at 0.95 | 65% |
| Recall@2G, random orderingwith complete-closure@2G of 1%; the doubled review budget alone hands out substantial partial credit | 0.54 |
| Complete recovery, Claude Sonnet 4.5, primary parserevery response arrived in a code fence; stripping the fence from the same responses gives 33% and Recall@2G 0.85 | 0% |
What this does not show
The answer key is the construction's own. Scores measure recovery of a generated dependency closure on machine-validated candidates, and the released protocol for human review, which checks each proposed dependency against the text and looks for affected statements the closure missed, has no reported results. Edits are numeric, and each item is a small, finite collection of constructed English documents derived from Wikipedia material. The review cutoff uses the hidden affected-set size; a deployed reviewer would have to choose a cutoff without the answer key. Four hosted models were added after the initial results were seen and ran with provider-default sampling, and no provider run was repeated to estimate generation variability, so the intervals reflect source sampling alone. The evaluation pools all 150 candidates where the original plan named the report partition alone, and the construction-control report records access to selection and report labels before baseline outputs were available. Affected downstream sentences open with their named entity, a possible cue the shortcut rankings do not test. Prior model exposure to the source passages is unassessed, and independent replay needs item-level responses and executable scoring code beyond the aggregate materials. Grok 4.6 returned five parsable responses, so its row says little about how it ranks.
The authors' abstract
Editing one fact can change a derived quantity, comparison, or other statement across linked documents. RippleKB evaluates whether systems can recover a constructed set of affected source units completely. Each item edits one source fact in a small set of linked documents. The proposed answer key is a hidden dependency closure, with unaffected units also present. A system returns a ranked list of source units, word for word, under a review rule tied to the constructed affected-set size. We regenerate every item from a public multi-hop corpus and test identifier-order, document-position, lexical, and supporting-fact rankings, alongside role and template concentration checks. We release the audit protocol under which the proposed closure is to be reviewed unit by unit. We report how far retrieval-style and language-model baselines get toward recovering the constructed closure completely on the machine-validated candidate set. Complete recovery tests whether a ranking includes the entire constructed target, while exact-copy measures separately assess preservation of source text.
More in Telegraph English
Cite
@misc{bei2026ripplekb,
title = {RippleKB: Finding What an Edit Changes Across Linked Documents},
author = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Shvets, Alexey and Kalaev, Dmitri},
year = {2026},
note = {ICLR 2027 submission},
url = {https://telegrapher.ai/research/ripplekb/}
}Builds on
- Trivedi et al. (2022). MuSiQue: Multihop Questions via Single-hop Question Composition.
- Goknil et al. (2016). A Rule-Based Change Impact Analysis Approach in Software Architecture for Requirements Changes.
- Cohen et al. (2024). Evaluating the Ripple Effects of Knowledge Editing in Language Models.
- Zhong et al. (2023). MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions.
- Logan IV et al. (2022). FRUIT: Faithfully Reflecting Updated Information in Text.