---
type: post
title: Finding most of what an edit changed is not finding all of it
date: '2026-10-09'
description: We scored whether rankings recover every statement an edit changes, not
  just most. Recall and completion disagree, sometimes at the edited sentence itself.
authors:
- Sisong Bei
- Mikhail L Arbuzov
- Ziwei Dong
- Alexey Shvets
- Dmitri Kalaev
paper: https://telegrapher.ai/research/ripplekb.md
html: https://telegrapher.ai/blog/ripplekb/
---

# Finding most of what an edit changed is not finding all of it

Someone corrects a stock count in one document. The sentence they touched is now right. In another document, a combined total built on that count is now wrong, and a later sentence saying combined stock sits below some threshold may have stopped being true. Search for the edit and you find one of three statements that need review: the one nobody was going to miss.

Retrieval is usually scored by how much of the relevant material comes back, on average. A reviewer after an edit needs to know something stricter: whether anything affected is still missing. In [the paper](/research/ripplekb/) we built a benchmark that scores that question directly. Its completion score orders systems differently from recall.

## The target includes the sentence you edited

The paper's conceptual example, abbreviated: the edit raises North's count from n to n + δ, enough for the combined total to reach the threshold q.

```
Edit:  North's count  n  →  n + δ

A  North stores n crates.              changes: the edited sentence
B  Combined stock is n + m crates.     changes: total = North + South
C  Combined stock is below q crates.   changes: the new total reaches q
D  South stores m crates.              unchanged, though it feeds the total
E  North opens at dawn.                unchanged, though it shares the subject
```

A, B and C need another look. D and E are there on purpose: E is about the same store, D feeds the very total that changes, and neither is affected. Being on topic, or being part of the arithmetic, is not the same as being changed.

RippleKB builds 150 items like this. Each starts from a MuSiQue record, whose questions compose facts across Wikipedia passages, and edits one numeric value in a source-grounded sentence. A generator writes the dependent statements and the rules connecting them, spread over three or four documents that hold 14 to 26 candidate units between them; four to seven of those units are affected. The dependency graph stays hidden, so the system has to work out the consequences from the edit and the text. Everything the graph reaches from the edit, the edited sentence included, is the item's *dependency closure*, which serves as the answer key.

The system ranks the candidates, giving each unit's identifier and its text copied byte for byte. The scorer reads down to a cutoff of twice the number of affected units, a number the system is not told, and takes two scores from that prefix. *Recall@2G* is the average fraction of affected units found. *Complete-closure@2G* is the share of items where all of them were found, which is what a reviewer working down to the cutoff needs.

## Checking the benchmark for shortcuts

Generated benchmarks can leak, and ours did. An earlier construction gave all 44 development items the same 14-unit layout with four ascending affected identifiers; sorting by identifier recovered every affected unit without reading a word. We withdrew it.

The rebuilt version assigns opaque identifiers after validation, shuffles document and unit order, and keeps tokens from the edited values out of downstream units. Five checks then run against it, each with pass criteria registered in advance. Four are shortcut rankings scored against the answer key; the fifth measures how concentrated the dependency shapes and affected positions are. The strongest ranking shortcut was lexical bridging, which orders units by word overlap or BM25 similarity to the edit, and it completed at most 20% of all 150 items. All five checks pass.

## Recall and completion order systems differently

All 150 items, primary parser, 95% intervals in parentheses; † marks models we added after seeing the initial results.

| System | Complete-closure@2G | Recall@2G |
|---|---|---|
| Random ordering | 1% (0–3) | 0.54 (0.51–0.57) |
| BM25 on the edit | 20% (14–27) | 0.66 (0.63–0.70) |
| BM25 + embedding rerank | 13% (8–19) | 0.68 (0.65–0.71) |
| Qwen3 32B | 53% (45–61) | 0.72 (0.66–0.78) |
| Llama 3.3 70B | 65% (57–73) | 0.91 (0.88–0.93) |
| Claude Sonnet 5 † | 82% (76–88) | 0.95 (0.93–0.97) |
| Claude Fable 5.1 † | 49% (41–57) | 0.89 (0.87–0.91) |
| Claude Opus 5 † | 0% (0–0) | 0.79 (0.78–0.81) |

Random ordering shows why completion needs a score of its own. With a budget twice the target size, an arbitrary order collects about half the affected units on average and finishes almost no items. The stronger language models lift both columns, Claude Sonnet 5 furthest. Even there, the distance between the columns is the point: a nearly complete set still leaves a statement nobody reviews.

The two scores can also move in opposite directions. Reranking BM25's candidates with Titan Text Embeddings V2 nudged the recall point estimate up and lowered complete recovery; a recall leaderboard would log a small gain, while a reviewer would get fewer edits whose consequences all landed in front of them. Treat that one as a hint. Both pairs of intervals overlap, and none of the three partitions the items were split into shows both moves at once. The firmer reversal comes from a model we added late.

### Opus 5 left the edited sentence below the cutoff

Claude Opus 5 beats BM25 on recall, with intervals that do not overlap, and completed none of the 150 items. Within its cutoff, its recovery of the directly edited sentence is zero: it ranked consequences and placed their source below the line. The edit record had already told it what changed, but that record does not stand in for the sentence in the ranked answer, and the edited sentence belongs to every item's impact set. Whatever else those rankings found, each of them was incomplete by construction.

Opus is not alone in its lean. Fable also recovers downstream units more often than the edited sentence; BM25 and Llama tilt the other way. We have not diagnosed why.

### Delivery and copying are separate losses

Our primary parser accepts the specified JSON object without repair, and an unreadable response becomes an empty ranking. Claude Sonnet 4.5 wrapped every response in a code fence and scored zero. Strip the fence from the same saved responses and it reaches 33% completion with recall 0.85. Qwen3 32B fenced 30 of its outputs, and those count as empty rankings too. The rule is harsh on purpose, since a pipeline that cannot read a ranking has no ranking; the reparse sits beside the primary score as a diagnosis, not a replacement.

Copying is scored apart from selection. Among the affected units a system did recover, exact copying is high for every system whose rankings parsed. Opus copies at 0.999, yet its numeric-span score, which also counts the units it missed, is 0.28.

## Report completion beside recall

If retrieval decides what gets reviewed after a change, report complete recovery at a stated review budget next to recall. The two can disagree in direction, not just size: clearly for Opus 5 against BM25, more softly for the embedding reranker. Unreadable output and altered copies are failures too, and each deserves its own count.

[Telegraph English](/research/telegraph-english/) rewrites text into lines that are each an independently addressable fact, and a store of addressable facts invites editing one fact at a time. Which other facts an edit made stale is the question RippleKB scores, though on sentences in constructed prose rather than TE lines: a shared problem, not a shared result. [What Survives Learned Symbolic Compression?](/research/what-survives-learned-symbolic-compression/) argues for the same habit of measuring separately, treating source fidelity, internal validity and what a consumer model recovers as different quantities. The empty-ranking rule has a counterpart in [Evaluating Relational Context Compression at Realized Token Budgets](/research/relational-context-compression/), which scores construction failures as empty answers.

## What is still open

Every score above measures recovery of a closure we constructed. We released a protocol for human review of that key: two reviewers recover affected units blind, then check the proposed closure unit by unit and edge by edge, then adjudicate. No review results are reported yet. A dependency the text does not support, or an affected statement the closure missed, would shift the scores by amounts we cannot size.

The setting is narrow: numeric edits in small collections of constructed English documents derived from Wikipedia material. The cutoff uses the constructed affected-set size, which isolates ranking completeness; a deployed reviewer would have to pick a cutoff without the answer key. One possible cue is untested, too. Affected downstream sentences open with their named entity, while unaffected ones have unconstrained openings, and none of the shortcut rankings probes that.

The models marked † and the fence-stripping reparse were both added after we saw the initial results, and the † models ran with provider-default sampling. No provider run was repeated to measure generation variability, so the intervals reflect source sampling alone. Before construction, source records were assigned to development, selection and a held-out report partition, the last holding 90 items; the original plan was to report on those 90 alone, and the scores here pool all 150. The construction record also shows the build process had access to the hidden labels of the selection and report partitions before any baseline output existed; a held-out partition would normally stay unseen until the end.

On these items, recall and completion order systems differently, and the difference can fall on the very sentence a reviewer would assume was covered. Whether the constructed closures survive human review is what the released protocol exists to find out.
