---
type: paper
slug: ripplekb
title: 'RippleKB: Finding What an Edit Changes Across Linked Documents'
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Alexey Shvets
- Dmitri Kalaev
date: '2026-08-25'
status: ICLR 2027 submission
line: Telegraph English
pages: 21
html: https://telegrapher.ai/research/ripplekb/
pdf: https://telegrapher.ai/papers/ripplekb/ripplekb.pdf
reader: https://telegrapher.ai/research/ripplekb/read/
json: https://telegrapher.ai/api/papers/ripplekb.json
openreview: https://openreview.net/forum?id=zRZ3MrSLIj
---

# RippleKB: Finding What an Edit Changes Across Linked Documents

## Paper gist

- **Claim:** A ranking can recover most of what an edit changed and still miss the edited sentence
- **TL;DR:** Asking whether a ranking finds every statement an edit changes, not just most, reorders systems. Embedding reranking nudged BM25's recall up and completed fewer sets; one model scored well on recall while missing the edited sentence itself.
- **Method:** 150 machine-validated items regenerated from MuSiQue, each a numeric edit over three or four linked documents with a hidden dependency closure of four to seven units, were ranked by deterministic baselines (random, render and identifier order, token overlap, BM25), BM25 with Titan Text Embeddings V2 reranking, Qwen3 32B, Llama 3.3 70B and Claude Sonnet 4.5, plus four hosted models added after the initial results were seen (Claude Sonnet 5, Claude Opus 5, Claude Fable 5.1, Grok 4.6), and scored by complete-closure@2G, Recall@2G, Recall@G and exact-copy fidelity under a no-repair parser.
- **Key result:** Complete recovery, Claude Opus 5: 0%; Complete recovery, BM25 with embedding reranking: 13%; Complete recovery, Llama 3.3 70B: 65%
- **Why it matters:** This is for people who maintain text that refers to other text: totals computed from figures held elsewhere, summaries that restate numbers from source pages, corpora that downstream answers lean on.
- **Limits:** The answer key is the construction's own.
- **Status:** ICLR 2027 submission, August 2026
- **Read:** reader /research/ripplekb/read/, PDF /papers/ripplekb/ripplekb.pdf

## Abstract

Editing one fact can change a derived quantity, comparison, or other statement across linked documents. RippleKB evaluates whether systems can recover a constructed set of affected source units completely. Each item edits one source fact in a small set of linked documents. The proposed answer key is a hidden dependency closure, with unaffected units also present. A system returns a ranked list of source units, word for word, under a review rule tied to the constructed affected-set size. We regenerate every item from a public multi-hop corpus and test identifier-order, document-position, lexical, and supporting-fact rankings, alongside role and template concentration checks. We release the audit protocol under which the proposed closure is to be reviewed unit by unit. We report how far retrieval-style and language-model baselines get toward recovering the constructed closure completely on the machine-validated candidate set. Complete recovery tests whether a ranking includes the entire constructed target, while exact-copy measures separately assess preservation of source text.

## What we did and found

In the paper's conceptual example, one sentence says the North store holds n crates, a second gives combined stock as n + m, and a third says combined stock is below q. Raise North's count far enough for the total to reach q and all three change. Two neighbours do not: South's count feeds the total but stays put, and a note that North opens at dawn shares the subject and stays true. The changed sentences form the item's impact set, which RippleKB proposes through a hidden dependency closure. Its 150 machine-validated items, regenerated from MuSiQue, each edit one numeric value across three or four linked documents of 14–26 candidate units, four to seven of them affected. A system sees the edit and the documents, then ranks units, giving each one's identifier and its text copied byte for byte. The scorer reads the top K = min(|U|, 2|G|) positions, a budget set by the hidden affected-set size, and reports two numbers from that one prefix: Recall@2G, the mean fraction found, and complete-closure@2G, the share of items found whole. Four shortcut rankings were tested against the key (identifier order, document position, lexical overlap with the edit, the original question's supporting facts), and lexical bridging came closest, completing at most 20% of the pooled items.

Recall and completion separate the systems, sometimes in opposite order. Under the doubled budget, random ordering already earns Recall@2G 0.54 while completing 1% of items. Reranking BM25 with embeddings raised the recall point estimate slightly and cut complete recovery from 20% to 13%. The stronger language-model rankings lift both: Llama 3.3 70B completed 65% of items at recall 0.91, and Claude Sonnet 5, one of four models added after the initial results were seen, 82% at 0.95. The gap is widest for Claude Opus 5, also a later addition, with recall 0.79 and no completed item; within its cutoff, it recovered the directly edited sentence on none of the items. Delivery is a separate loss. Claude Sonnet 4.5 wrapped every response in a code fence and scored zero under the no-repair parser, while the same saved responses, stripped of the fence, give 33% completion and recall 0.85. Copying is a third. Conditional exactness counts recovered units alone and runs high, from 0.942 for Llama upward among systems with parsed rankings. Opus's 0.999 on that measure sits beside numeric-span preservation of 0.28, because the span score also counts the units it left out.

## Key numbers

| Measure | Value |
|---|---|
| Complete recovery, Claude Opus 5 | 0% |
| Complete recovery, BM25 with embedding reranking | 13% |
| Complete recovery, Llama 3.3 70B | 65% |
| Recall@2G, random ordering | 0.54 |
| Complete recovery, Claude Sonnet 4.5, primary parser | 0% |

## Why it matters

This is for people who maintain text that refers to other text: totals computed from figures held elsewhere, summaries that restate numbers from source pages, corpora that downstream answers lean on. After a correction, the operational question is whether anything affected is still unreviewed. Average recall does not answer it. A reranker that edges recall up can finish fewer review sets, and a model can post a respectable recall while leaving the very sentence that was edited outside its review cutoff. Report complete recovery beside recall, at a stated review budget, and both failures show.

The paper also keeps apart failures that a single score would merge. An unparsable response counts as an empty ranking, so delivery is part of what is measured; harsh, but a pipeline that cannot read a ranking has no ranking. The fence-stripping reparse of the same responses is reported beside it as a diagnosis, not a replacement. Selection and copying are scored apart too, since a correct identifier can carry altered text and perfect copies can still leave members of the set missing.

RippleKB sits with the Telegraph English papers through its unit of account. Telegraph English rewrites text into lines that are each an independently addressable fact. A store of addressable facts invites edits one fact at a time, and RippleKB scores the question that follows: which other units did the edit change? Its units are sentences in constructed prose documents rather than TE lines, so the link is a shared question, not a shared result. The habit of measuring separately also runs through What Survives Learned Symbolic Compression?, which treats source fidelity, internal validity and what a consumer model recovers as different quantities.

## What this does not show

The answer key is the construction's own. Scores measure recovery of a generated dependency closure on machine-validated candidates, and the released protocol for human review, which checks each proposed dependency against the text and looks for affected statements the closure missed, has no reported results. Edits are numeric, and each item is a small, finite collection of constructed English documents derived from Wikipedia material. The review cutoff uses the hidden affected-set size; a deployed reviewer would have to choose a cutoff without the answer key. Four hosted models were added after the initial results were seen and ran with provider-default sampling, and no provider run was repeated to estimate generation variability, so the intervals reflect source sampling alone. The evaluation pools all 150 candidates where the original plan named the report partition alone, and the construction-control report records access to selection and report labels before baseline outputs were available. Affected downstream sentences open with their named entity, a possible cue the shortcut rankings do not test. Prior model exposure to the source passages is unassessed, and independent replay needs item-level responses and executable scoring code beyond the aggregate materials. Grok 4.6 returned five parsable responses, so its row says little about how it ranks.

## Blog post

[Finding most of what an edit changed is not finding all of it](https://telegrapher.ai/blog/ripplekb.md)

## Cite

```bibtex
@misc{bei2026ripplekb,
  title         = {RippleKB: Finding What an Edit Changes Across Linked Documents},
  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Shvets, Alexey and Kalaev, Dmitri},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/ripplekb/}
}
```
