RippleKB: Finding What an Edit Changes Across Linked Documents
Back to the paper page. ICLR 2027 submission, August 2026.
All 21 pages are shown below.
Text of page 1
R IPPLE KB: F INDING W HAT A CROSS L INKED D OCUMENTS A BSTRACT Editing one fact can change a derived quantity, comparison, or other statement across linked documents. RippleKB evaluates whether systems can recover a constructed set of affected source units completely. Each item edits one source fact in a small set of linked documents. The proposed answer key is a hidden dependency closure, with unaffected units also present. A system returns a ranked list of source units, word for word, under a review rule tied to the constructed affected-set size. We regenerate every item from a public multi-hop corpus and test identifier-order, document-position, lexical, and supporting-fact rankings, alongside role and template concentration checks. We release the audit protocol under which the proposed closure is to be reviewed unit by unit. We report how far retrieval-style and language-model baselines get toward recovering the constructed closure completely on the machine-validated candidate set. Complete recovery tests whether a ranking includes the entire constructed target, while exact-copy measures separately assess preservation of source text. E DIT C HANGES Anonymous authors Paper under double-blind review AN 1 I NTRODUCTION An edit can be locally correct while leaving a collection of documents inconsistent. A changed quantity can alter a total in another document and an explanation farther along the chain. Finding the edited sentence alone leaves those consequences unreviewed. Change-impact analysis studies dependencies in requirements and software artifacts; document-update research studies revisions prompted by new evidence (Goknil et al., 2016; Etezadi et al., 2025; Logan IV et al., 2022). The difficult endpoint is finding every affected statement. A system can retrieve useful evidence or produce a correct revised answer while missing statements outside that answer’s evidence path. Document editing, knowledge editing, and multi-hop question answering evaluate related capabilities through revised text, changed model knowledge, or evidence for a chosen question (Kruthof, 2026; Dwivedi-Yu et al., 2024; Cohen et al., 2024; Zhong et al., 2023; Yang et al., 2018). RippleKB makes the affected source units themselves the target. The target is an impact set: the source units whose truth or content changes because of the edit. RippleKB proposes it through a dependency closure, the set reached through an item’s hidden proposed dependencies. The system ranks each verbatim unit, a source unit reproduced exactly as it appears in its document. Unaffected statements remain among the candidates, so the ranking determines where the complete constructed target appears. Figure 1 follows a conceptual stock-count edit through a changed total and comparison, then shows an affected unit left beyond the ranking’s review cutoff. Partial recovery and complete recovery therefore answer different maintenance questions. One measures the fraction of the impact set found; the other asks whether any affected unit remains missing. The scorer applies the same size-based budget rule to each ranking, using an affected-set size hidden from the system. Selection and copying are also distinct: a correct identifier can retrieve a unit even when its submitted text is altered. The construction makes the proposed answer key explicit and tests specified structural and lexical routes to recovering it. We evaluate retrieval and language-model rankings on the resulting machinevalidated candidates. Embedding reranking raises BM25’s partial-recall point estimate from 0.66 to 0.68 while lowering complete recovery from 20% to 13%. A later-added language model reaches 1 Reviewers: please read the Reviewer Guidelines (iclr.cc/Conferences/2027/ReviewerGuidelines) and the AI Policy for Reviewers (iclr.cc/Conferences/2027/AIPolicyForReviewers). If you used AI to expand, edit, or polish your review, please provide the input text to the LLM. Better still, consider skipping the LLM and submitting your original text: we, and the authors, are much more interested in your unedited thoughts than in what an LLM has to say. AI-assisted or not, you are putting your name and reputation behind your review: LLM-generated falsehoods, hallucinations or misrepresentations are subject to disciplinary action, which may include desk-rejecting all papers you have authored.
Text of page 2
An edit propagates through linked statements
Edit record: A: n → n + δ
n, m, q are fixed pre-edit values; δ > 0.
Assume n + m < q ≤ n + δ + m so the comparison changes.
Derived total B
Comparison C
North stores n crates.
Combined stock is
n + m crates.
Combined stock is
below q crates.
After: n + δ
After: n + δ + m
After: not below q
Rules: total = North + South; comparison = (total < q).
Unchanged: D South stores m crates. E North opens at dawn.
Source unit A
Illustrative ranking
Coverage and copying differ
A North stores n crates.
B Combined stock is n + m crates.
… other unchanged units …
A and B are copied exactly.
C falls below the cutoff.
Full coverage needs A, B, C;
returned words stay unchanged.
scorer's top K cutoff
C Combined stock is below q crates.
G = {A, B, C}; K = min(|U|, 2|G|). The edit record is excluded.
Conceptual example; lists are abbreviated and the ranking is illustrative.
Figure 1: When does a useful ranking still miss the full impact? A conceptual count edit changes a
total and comparison; the illustrative ranking leaves an affected unit beyond its review cutoff.
0.79 recall but zero completion because it misses the directly edited sentence within the scored prefix.
These cases make complete recovery an observable endpoint beyond average recall; parsing and exact
copying then determine what source evidence the response delivers.
RippleKB supplies a controlled construction, a complete-recovery task, and paired measurements of
selection and source copying.
2
T HE TASK : RECOVER THE AFFECTED SOURCE UNITS
An item couples a visible edit and linked documents with a hidden proposed answer key. The edit
names a source sentence and its original and replacement values. The documents contain candidate
units, each identified by a stable opaque identifier. The system receives the edit, the documents, the
candidate identifiers, and their text. The dependency graph, impact-set size, source-record identifier,
and split assignment remain hidden during prediction. The system infers the affected statements from
their text and returns them for review.
The constructed dependency closure specifies the evaluation target; the released human-review
protocol specifies how to assess whether that membership matches the text. Candidate items contain
three or four documents and 14–26 units, with proposed target sets of four to seven units and
maximum depths of three to five. The separately displayed edit record is outside the impact set; the
edited source sentence belongs to it. Depth counts the edge from that record to the source sentence,
so the edited sentence is direct and its consequences are downstream. Units outside the proposed
closure remain in the documents, making the target a subset of the available text.
Following an edit through the example. The conceptual stock example in Figure 1 starts with
a local count, then follows its contribution to a combined total and a comparison against a fixed
threshold. Changing the count changes the total because the document supplies the addition rule.
Changing the comparison additionally requires the total to cross the stated threshold; the figure gives
2
Text of page 3
that condition explicitly. The dependency is therefore about the consequence of this edit under the
stated rule, rather than the mere occurrence of a related quantity.
The unchanged statements in the example make this distinction visible. A statement about when the
same store opens shares its subject with the edit but is unaffected by the count change. The other
store’s count contributes to the total while remaining unchanged itself. The target follows the edit’s
consequences, so topical similarity and participation in the surrounding calculation are different
reasons for a unit to appear in the documents. The hidden graph records the proposed affected set;
the visible text supplies what the system can use to recover it.
Figure 1 illustrates the output contract; Section 3 gives the numeric-edit and lexical constraints used
for actual candidates. Returning a verbatim unit means copying its supplied text, not rewriting it to
incorporate the new value. This keeps finding affected material separate from deciding how to revise
it. The output contract asks the system to rank every candidate unit; the scorer also accepts shorter
lists, with missing positions receiving no credit. Each entry contains its identifier and source text. An
unknown or duplicate identifier consumes a position without adding selection credit. A valid affected
identifier can receive selection credit even when its text is copied incorrectly; fidelity scores measure
that separate error.
The review budget is an item-specific scoring allowance, determined by a common rule. For
candidate set U and target set G from the constructed dependency closure, the larger budget is
K = min(|U |, 2|G|). The scorer knows |G|; the system ranks candidates without seeing that size or
its resulting cutoff. Returning every candidate therefore earns credit according to its order within the
same budget. Appendix E gives the duplicate, malformed-output, and missing-position rules.
Partial recovery and completed review sets. Three retrieval measures describe the same ranking
at complementary resolutions. Recall@2G is the mean fraction of affected identifiers recovered
within K positions; complete-closure@2G is the proportion of items whose entire impact set is
recovered there. Recall@G uses the smaller budget min(|U |, |G|). The recall average gives each
item equal weight, regardless of the size of its impact set. Complete recovery instead records an
all-or-nothing outcome for each item. Writing P i (k) for the distinct valid identifiers recovered in the
first k positions makes the two per-item scores explicit:
r i =
|P i (K i ) ∩ G i |
,
|G i |
c i = 1{G i ⊆ P i (K i )},
K i = min(|U i |, 2|G i |).
The reported recall and completion scores average r i and c i across items. The stricter prefix asks
which units the ranking prioritizes; the larger prefix gives room for unaffected units while still
requiring all affected identifiers for completion.
The larger budget gives random selection substantial partial credit on the current candidate collection.
The retained random baseline reaches Recall@2G 0.54 alongside complete-closure@2G 1%, with
Recall@G 0.27. These are the pooled baseline scores used in Section 5.
Finding a unit and preserving its bytes. Canonical comparison converts transport CRLF line
endings to LF and otherwise requires exact source bytes. Numeric and named-entity fidelity likewise count preserved annotated spans, with missed affected units contributing no recovered spans.
Conditional verbatim exactness considers only recovered affected identifiers and pools those records
across items. Numeric and entity scores pool their respective annotated spans rather than averaging
item-level fractions. Each span must retain its bytes at the recorded position in the returned unit;
reformatting a number or shifting a span can fail this test without changing its meaning. Appendix E
gives the remaining ranking and fidelity definitions.
Table 1 collects the questions answered by these measurements.
3
C ONSTRUCTION AND ANSWER - KEY REVIEW
Construction must supply a target that is explicit enough to score and whose proposed dependencies
can be assessed against the documents. Mechanical consistency, shortcut controls, and semantic
review address different parts of that requirement. The collection contains 150 machine-validated
candidate items.
3
Text of page 4
Table 1: What each measurement asks about the returned ranking. Set recovery, complete review, and source preservation retain their different denominators. Measurement Question and aggregation Recall@2G How much of each target was found? Mean of item fractions. Was the whole target found? Fraction of completed items. Was recovered text copied exactly? Pools recovered affected units. Were annotated bytes preserved? Pools all affected numeric or entity spans. Complete-closure@2G Conditional exactness Span fidelity 3.1 S OURCE ASSIGNMENT AND CANDIDATE GENERATION Each item is regenerated from MuSiQue, whose questions compose facts across Wikipedia passages (Trivedi et al., 2022). Public source identifiers are assigned to development, selection, and held-out report partitions before construction. Source records used by another in-review work on the same corpus were excluded before assignment. Appendix K records source isolation. The generation process fixes the edit, dependency shape, and document placement before producing downstream text. The edit changes a numeric value in a source-grounded sentence. A generator supplies proposed dependent statements and the relation text needed to connect them. Eight dependency shapes vary the impact-set size and depth, while unaffected units provide alternatives within every document. Each proposed dependency closure crosses a document boundary and contains mostly units outside the source question’s supporting facts. Graph reachability proposes which units change; the documents must still support that interpretation. A statement that eligibility depends on a count, for example, needs a rule connecting the particular count change to a different eligibility outcome. The edge records the intended dependence; the released review protocol specifies how to assess it against the rendered statements. Mechanical checks enforce the construction’s explicit contract. They verify required units and edges, source grounding, placement, unique unit text, and the alignment of annotated spans. Affected downstream units begin with their annotated entity, while unaffected units use unconstrained sentence openings. The checks also prevent tokens from the edited values from appearing in titles and downstream units, limiting a direct lexical cue. Public identifiers are assigned independently of construction roles, and document and unit orders are shuffled. The first mechanically valid proposal in each predetermined source reserve is selected (Appendix B). 3.2 T ESTING ROUTES TO THE ANSWER KEY Identifier ordering tests both sort directions; document-position ranking learns affected positions from development items and applies that ordering to the other partitions. Lexical bridging ranks units by BM25 or token overlap with the edit. The supporting-fact control places the original question’s support units first, followed by a seeded ordering of the remaining units. It uses hidden support labels as a privileged diagnostic of how far that evidence alone reaches. Evidence sufficient for that question may cover only part of the impact set. Role and template concentration is a composition check rather than a ranking method. It examines dependency-shape coverage, affected-position patterns, diversity, and uniqueness. Figure 2 separates these composition measurements from rankings; Appendix C gives the five checks’ thresholds and full results. Lexical bridging is the strongest tested shortcut, completing at most 20% of pooled candidates and 27% of the selection split. All five checks meet their registered criteria; the entity-first openings of affected downstream units remain a potential cue outside these rankings. The plotted lexical and identifier summaries take separate maxima for each metric. Appendix C gives split coverage, thresholds, and the distinct random orderings used for construction diagnostics and baseline evaluation. 4
Text of page 5
Can simple structure identify the affected units? Construction checks against the constructed dependency closure Lexical selection: complete 26.7% (BM25); recall 0.64. Ranking family / split Identifier order* Recall@2G 0 0.5 All candidates Development Selection Report Document position Selection Report Lexical bridging* All candidates Development Selection Report Supporting facts All candidates Development Selection Report Observed Complete-closure@2G Registered limit 0.2 0.4 0.6 0.6 0.7 0.8 0.9 Construction random (pooled) * Per-metric maxima can come from different strategies. Composition: max. shape 15.3%; affected-position pattern 1.3% (not retrieval rates). Figure 2: Can the specified shortcuts recover the whole closure? Ranking controls pair complete recovery with recall; the separate composition checks measure concentration of dependency shapes and affected positions. 3.3 C ONSTRUCTED TARGETS AND TEXTUAL SUPPORT The reported scores measure recovery of the construction-defined closure on machine-validated candidates. The released review protocol assesses whether the visible text supports each proposed dependency and whether other affected units are missing from the closure (Appendix D). 3.4 C ANDIDATE COMPOSITION The construction assigns 30 candidates to development, 30 to selection, and 90 to the report partition. The present evaluation pools all 150 candidates without tuning on any split; Appendix G retains the available partition-level scores. Appendix B gives the dependency-shape, document-count, impact-size, and depth composition. 4 E VALUATION PROTOCOL The comparison measures how early systems rank a complete impact set on the same 150 machinevalidated candidates. The initial configurations were fixed before candidate scoring, and no setting was tuned on the development, selection, or report partition. The pooled evaluation uses all construction partitions, whereas the original plan specified a report-only population. Appendix A gives the source-cluster uncertainty procedure and Appendix F records the retained runs. The initial comparison spans deterministic rankings, retrieval with embedding reranking, and language models. Deterministic methods use random, rendered, or identifier order, token overlap, or BM25 5
Text of page 6
relevance to the edit. The retrieval baseline reranks BM25 candidates using Titan Text Embeddings V2. Qwen3 32B, Llama 3.3 70B, and Claude Sonnet 4.5 receive the edit and documents and rank candidate identifiers with each verbatim unit’s original text. Claude Sonnet 5, Claude Opus 5, Claude Fable 5.1, and Grok 4.6 were added after the first results were seen. Their displays identify this later group explicitly. They use the same prompt, output allowance, and scoring contract, but reject the specified temperature field and run with providerdefault sampling. Fable also rejects the thinking-disabled setting; Appendix F records the effective request for each family. The primary parser accepts the specified JSON object without response repair. An invalid response receives an empty ranking, so formatting failures remain part of the evaluated pipeline. A separately labelled format-tolerant parse strips one surrounding code fence from the same recorded responses. It diagnoses a transport-format effect without replacing the primary scores or generating new responses. Uncertainty intervals resample public source identifiers, retaining the items associated with each sampled source and the multiplicity of repeated draws. The displayed 95% intervals use 10,000 percentile-bootstrap replicates of the item-weighted statistic. They describe source-sampling variation for the retained runs; provider runs were not repeated to estimate sampling variability from generation. 5 R ECOVERING THE WHOLE IMPACT SET The paired scores reveal which systems assemble complete targets and which mainly recover parts of them. Figure 3 places complete-closure@2G beside Recall@2G for every retained system on the machine-validated candidate set. Both use the same reviewed prefix of each ranking, making the difference a property of the recovered set rather than the review allowance. 5.1 P ARTIAL RECOVERY AND COMPLETION SEPARATE THE SYSTEMS Llama 3.3 70B reaches complete-closure@2G 65% with Recall@2G 0.91, compared with BM25’s 20% and 0.66. The later-added Sonnet 5 reaches 82% and 0.95, the highest retained completion estimate. Thus the stronger language-model rankings improve both the amount of affected material found and the frequency of finishing the target. Their remaining gap between recall and completion identifies the endpoint that RippleKB measures: a nearly complete set still leaves a statement unreviewed. Retrieval refinements illustrate why both scores are useful. BM25 with embedding reranking reaches 13% complete recovery and Recall@2G 0.68, compared with BM25’s 20% and 0.66. A higher partial-recall point estimate can accompany fewer completed targets. The stricter review allowance tests whether affected units appear early. Recall@G is 0.37 for BM25, 0.78 for Llama, and 0.89 for the later-added Sonnet 5. These scores complement the paired largerbudget results above by showing how much target content occupies the first target-sized prefix. Appendix G retains all scores, intervals, and partition-level summaries. 5.2 T HE MISSING UNIT CAN BE THE DIRECTLY EDITED SENTENCE The depth summaries make the complete-recovery endpoint more concrete. BM25 and Llama recover the directly edited unit more often than downstream units, whereas the later-added Opus and Fable show the reverse ordering. Opus reaches Recall@2G 0.79 with complete-closure@2G 0%; its directunit recovery within the review prefix is zero. Its partial score reflects recovery elsewhere in the proposed dependency closure, while the absent direct unit prevents completion. This distinction matters because the task includes the source sentence as well as consequences of its edit. The separate edit record tells the system what changes; it does not replace that sentence in the ranked answer. A ranking of downstream consequences alone can therefore be useful yet incomplete under the same target definition. Appendix J describes the supplied direct and downstream summaries and their scope. 6
Text of page 7
Recovering many units can still leave the closure incomplete All 150 candidates; primary parser; source-cluster intervals System Complete-closure@2G Recall@2G Random ordering Render order Identifier order (ascending) Identifier order (descending) Token overlap with edit BM25 on edit BM25 + embedding rerank Qwen3 32B Llama 3.3 70B Claude Sonnet 4.5 (parse failure) Hosted additions: added after the first results were seen These four models used provider-default temperature. Claude Sonnet 5 Claude Opus 5 Claude Fable 5.1 Grok 4.6 0 50 100% 0 0.5 1.0 Figure 3: Complete recovery and recall under the primary response contract, with 95% source-cluster intervals. Sonnet 4.5’s zero scores reflect fenced responses; Opus’s zero completion reflects omission of the directly edited unit from the review prefix; Grok supplies only five parsed responses. The lower group contains models added after the first results were seen. Appendix H reports fence-removal rescoring of the same responses. 5.3 R ESPONSE DELIVERY DETERMINES THE PRIMARY RANKING The primary scores evaluate the delivered JSON response under the required parser. Sonnet 4.5 fences every response and therefore receives complete-closure@2G 0% with Recall@2G 0.00. The separately reported fence-stripping parse of those same responses gives 33% and 0.85. Qwen’s fenced responses similarly contribute failures under primary scoring; Appendix H places the two parses side by side. This secondary measurement isolates a response-wrapper effect while holding the recorded model text fixed. It helps distinguish an unavailable ranking from a parsed ranking that omits affected units. The later-added Grok run supplies only five parsed responses under the 4,096-token output cap, reaching 3% complete recovery with Recall@2G 0.03. Empty and malformed responses remain in the full denominator. Table 2: Removing a response wrapper changes access to the ranking. These are the two families with fenced responses; all scores retain 150 items. Only the surrounding fence is removed. Model Parse Parsed Complete@2G Recall@2G Qwen3 32B Primary Fence stripped 120 150 53% 66% 0.72 0.90 Sonnet 4.5 Primary Fence stripped 0 150 0% 33% 0.00 0.85 7
Text of page 8
Table 2 shows the same-response comparison directly. Qwen’s 30 fenced outputs become available to the scorer, while stripping Sonnet’s fences reveals a ranking with high partial recall and substantially fewer complete sets. 5.4 Retrieval controls reproduce selected units exactly, yet their complete-recovery scores remain low. Among recovered affected units, Llama’s conditional verbatim exactness is 0.942 and Qwen’s is 0.995. The later-added Sonnet 5 reaches 0.998. These copying scores condition on finding an affected unit: a correctly selected identifier can carry altered text, while perfectly copied units can still leave other members of the impact set missing. Numeric and named-entity fidelity include annotated spans in missing units in their denominators. Those measures therefore combine coverage with preservation, while conditional exactness measures copying after recovery. Appendix I reports them together with their respective denominators. Table 3: Copying and coverage can diverge. Exact copying conditions on recovered affected units; the span columns retain all annotated affected spans, including missing units. Selected systems discussed in the text are shown; Appendix I gives every system. System C OPYING FIDELITY ANSWERS A SEPARATE QUESTION BM25 Qwen3 32B Llama 3.3 70B Exact copying Numeric spans Entity spans 1.000 0.995 0.942 0.91 0.75 0.97 0.67 0.72 0.91 0.94 0.28 0.95 0.79 Added after the first results were seen Sonnet 5 0.998 Opus 5 0.999 Opus’s near-perfect conditional copying coexists with numeric-span preservation of 0.28 (Table 3). The conditional score describes its recovered units; the broader denominator also includes the annotated material it failed to return. 6 R ELATED WORK Finding affected artifacts. Change-impact analysis provides the closest task connection. Requirement analysis propagates relations among linked artifacts, while ProReFiCIA ranks requirements affected by a change rationale against expert labels (Goknil et al., 2016; Etezadi et al., 2025). RAID propagates edits through a compositional knowledge base and evaluates affected identifiers (Guo et al., 2026). RippleKB examines affected-unit recovery in linked free-form text, with unchanged source wording included in the output. The review allowance makes the ordering of affected and unaffected units part of the task. The target is the complete constructed set presented for review. Updating text and model knowledge. FRUIT evaluates faithful Wikipedia updates, EditEval studies instruction-guided editing, and EditPropBench distinguishes required revisions from protected text (Logan IV et al., 2022; Dwivedi-Yu et al., 2024; Kruthof, 2026). FRESCO examines retrieval as corpora change (An et al., 2026). RippleKB evaluates the source statements selected for possible revision, before a system rewrites them. Knowledge-editing benchmarks instead examine consequences within a model: RippleEdits studies related changes, MQuAKE uses multi-hop questions, and logical-rule evaluations probe derived consequences (Cohen et al., 2024; Zhong et al., 2023; Moteu Ngoli et al., 2025). Locality evaluations and RippleBench examine unaffected cases and indirect effects of unlearning (Singh, 2026; Rinberg et al., 2025). Our affected and unaffected units remain jointly visible in the documents being ranked. Constructing dependency-sensitive evaluations. Analyses of HotpotQA identify single-passage and disconnected-reasoning routes to correct answers (Chen & Durrett, 2019; Min et al., 2019; Trivedi et al., 2020). MuSiQue composes questions to require connected reasoning and supplies the public source records used here (Trivedi et al., 2022). Its constituent questions expose the intended reasoning 8
Text of page 9
structure, while separate artifact-based models test how much of the benchmark can be solved through other routes. These studies motivate examining whether a dataset’s visible regularities reveal its labels. Such diagnostics matter when generation can make affected text look different from the unaffected alternatives. RippleKB tests identifier order, document positions, lexical overlap, original supporting facts, and composition regularity against its constructed targets. The released closure-review protocol separately asks whether the generated text supports the proposed dependencies and whether affected statements are missing. Selecting evidence and preserving text. BERTSUM learns extractive selection; RECOMP, FILCO, and Provence select or compress context for subsequent model use (Liu & Lapata, 2019; Xu et al., 2024; Wang et al., 2023; Chirkova et al., 2025). Symbolic rewriting produces compact fact lines, and matched-budget studies compare those representations with prose summaries in multi-hop question answering (Arbuzov et al., 2026; Bei et al., 2026). Cross-document evidence evaluation likewise studies responses supported by distributed information (Eletter et al., 2026). RippleKB’s exact-copy requirement retains each selected unit as an inspectable source object. Its scores distinguish selecting the right identifiers, completing the target, and preserving the returned text. 7 D ISCUSSION RippleKB makes a concrete retrieval failure visible: high partial recall can coexist with an unfinished impact set, including a ranking that misses the edited sentence itself. Make completion an explicit objective. A ranking can be useful before it contains the complete impact set. The paired recovery measures show how much of a constructed target has been found and whether any member remains beyond the reviewed prefix. Reporting both connects an average retrieval score to the maintenance question that motivates the task. The direct and downstream summaries sharpen this distinction: missing the edited sentence can prevent completion even when many of its proposed consequences are recovered. Review budgets therefore belong alongside recovery scores when comparing systems. Separate selection, delivery, and copying. The response contract creates an observable path from a model’s output to the source units a reviewer receives. The primary parser measures delivery under that contract. After a response is parsed, identifier recovery and exact copying answer separate questions. Conditional copying and selection-sensitive span preservation identify these different sources of loss. Match construction checks to the claim. The proposed dependency graph makes the target inspectable, and the named controls test specific ways to recover it from visible regularities. Whether that key matches all statements changed by the edit remains a semantic question; the released human-review protocol specifies how to assess it. RippleKB currently studies numeric edits in finite collections of constructed English documents derived from Wikipedia material. Within that setting, it exposes a review task that is easy to hide behind partial success: assembling every member of a proposed affected set and returning its source text intact. Its gold-size scoring allowance isolates ranking completeness; deploying the task in a review workflow would also require choosing a cutoff without access to the answer key. 9
Text of page 10
R EPRODUCIBILITY STATEMENT The materials provide construction and metric specifications, aggregate result displays, responseconfiguration records, and hashes of referenced inputs. The figure scripts and fixed data extracts rebuild the manuscript displays. Appendix K identifies the item-level artifacts and executable components required for full replay. Timestamps are withheld from the anonymous version and restored at camera-ready; Appendix L retains the ordered changes. E THICS STATEMENT The construction uses public source material under the documented upstream license and adds generated linked statements for evaluation. Redistribution requires source attribution, license notices, and a description of those modifications. The constructed documents are evaluation items rather than updates to the source encyclopedia. AI U SE S TATEMENT Generative models produce linked statements during benchmark construction, as described in Section 3. In this work, generative AI tools were used to polish the manuscript prose and assist with preparing explanatory illustrations. Generative AI tools were not used to design the model. All AI-assisted text and visual materials were reviewed and edited by the authors. The authors take responsibility for the final content of this paper, including all text, claims, and artifacts produced with AI assistance. R EFERENCES Sohyun An, Hayeon Lee, Shuibenyang Yuan, Chun-cheng Jason Chen, Cho-Jui Hsieh, Vijai Mohan, and Alexander Min. FRESCO: Benchmarking and Optimizing Re-rankers for Evolving Semantic Conflict in Retrieval-Augmented Generation. arXiv preprint arXiv:2604.14227, 2026. URL https://arxiv.org/abs/2604.14227. Mikhail L. Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, and Alexey A. Shvets. Telegraph English: Semantic prompt compression via structured symbolic rewriting, 2026. URL https: //arxiv.org/abs/2605.04426v1. Sisong Bei, Mikhail L. Arbuzov, Ziwei Dong, Dmitri Kalaev, and Alexey Shvets. Context compression is not one thing: Readable symbolic re-expression vs. coherent summary at matched budget, 2026. URL https://arxiv.org/abs/2606.14875v1. A. Colin Cameron, Jonah B. Gelbach, and Douglas L. Miller. Bootstrap-Based Improvements for Inference with Clustered Errors. Review of Economics and Statistics, 90(3):414–427, 2008. doi: 10.1162/rest.90.3.414. URL https://doi.org/10.1162/rest.90.3.414. Jifan Chen and Greg Durrett. Understanding Dataset Design Choices for Multi-hop Reasoning. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4026–4032, 2019. doi: 10.18653/v1/N19-1405. URL https://aclanthology.org/N 19-1405/. Nadezhda Chirkova, Thibault Formal, Vassilina Nikoulina, and Stéphane Clinchant. Provence: efficient and robust context pruning for retrieval-augmented generation. In International Conference on Learning Representations, 2025. URL https://proceedings.iclr.cc/paper_fi les/paper/2025/hash/5e956fef0946dc1e39760f94b78045fe-Abstract-C onference.html. Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. Evaluating the Ripple Effects of Knowledge Editing in Language Models. Transactions of the Association for Computational Linguistics, 12:283–298, 2024. doi: 10.1162/tacl_a_00644. URL https://aclanthology .org/2024.tacl-1.16/. 10
Text of page 11
Jane Dwivedi-Yu, Timo Schick, Zhengbao Jiang, Maria Lomeli, Patrick Lewis, Gautier Izacard, Edouard Grave, Sebastian Riedel, and Fabio Petroni. EditEval: An Instruction-Based Benchmark for Text Improvements. In Proceedings of the 28th Conference on Computational Natural Language Learning, pp. 69–83, 2024. doi: 10.18653/v1/2024.conll-1.7. URL https://aclanthology .org/2024.conll-1.7/. Saadeldine Eletter, Ruihong Zeng, Yuxia Wang, Maxim Panov, Aleksandr Rubashevskii, and Preslav Nakov. MIRAGE: Defending Long-Form RAG Against Misinformation Pollution. arXiv preprint arXiv:2607.05069, 2026. URL https://arxiv.org/abs/2607.05069. Romina Etezadi, Sallam Abualhaija, Chetan Arora, and Lionel Briand. LLM-Driven Cost-Effective Requirements Change Impact Analysis. arXiv preprint arXiv:2511.00262, 2025. URL https: //arxiv.org/abs/2511.00262. Arda Goknil, Ivan Kurtev, and Klaas van den Berg. A Rule-Based Change Impact Analysis Approach in Software Architecture for Requirements Changes. arXiv preprint arXiv:1608.02757, 2016. URL https://arxiv.org/abs/1608.02757. Jiajing Guo, Xueming Li, Jorge Piazentin Ono, Wenbin He, and Liu Ren. Scaling Expert Feedback with Reflective Edit Propagation in Compositional Knowledge Bases. In Proceedings of the ACM Conference on AI and Agentic Systems, 2026. doi: 10.1145/3786335.3813201. URL https://arxiv.org/abs/2606.05023. Garvin Kruthof. EditPropBench: Measuring Factual Edit Propagation in Scientific Manuscripts. arXiv preprint arXiv:2605.02083, 2026. URL https://arxiv.org/abs/2605.02083. Yang Liu and Mirella Lapata. Text Summarization with Pretrained Encoders. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3730–3740, 2019. doi: 10.18653/v1/D19-1387. URL https://aclanthology.org/D19-1387/. Robert L. Logan IV, Alexandre Passos, Sameer Singh, and Ming-Wei Chang. FRUIT: Faithfully Reflecting Updated Information in Text. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 3670–3686, 2022. doi: 10.18653/v1/2022.naacl-main.269. URL https://aclanthology.org/2022.naacl-main.269/. Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, and Luke Zettlemoyer. Compositional Questions Do Not Necessitate Multi-hop Reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4249–4257, 2019. doi: 10.18653/v1/P19-1416. URL https://aclanthology.org/P19-1416/. Tatiana Moteu Ngoli, N’Dah Jean Kouagou, Hamada M. Zahera, and Axel-Cyrille Ngonga Ngomo. Benchmarking Knowledge Editing using Logical Rules. In The Semantic Web – ISWC 2025, volume 16141 of Lecture Notes in Computer Science, pp. 41–56. Springer, 2025. doi: 10.1007/ 978-3-032-09530-5_3. URL https://link.springer.com/chapter/10.1007/97 8-3-032-09530-5_3. Roy Rinberg, Usha Bhalla, Igor Shilov, Flavio P. Calmon, and Rohit Gandikota. RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories. arXiv preprint arXiv:2512.04144, 2025. URL https://arxiv.org/abs/2512.04144. Stephen Robertson and Hugo Zaragoza. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 2009. doi: 10.1561/1500000019. URL https://doi.org/10.1561/1500000019. Aditya Pratap Singh. On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study. arXiv preprint arXiv:2608.26292, 2026. URL https://arxiv.org/abs/2608.26292. Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. Is Multihop QA in DiRe Condition? Measuring and Reducing Disconnected Reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 8846–8863, 11
Text of page 12
2020. doi: 10.18653/v1/2020.emnlp-main.712. URL https://aclanthology.org/2020. emnlp-main.712/. Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. MuSiQue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics, 10:539–554, 2022. doi: 10.1162/tacl_a_00475. URL https: //aclanthology.org/2022.tacl-1.31/. Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. Learning to Filter Context for Retrieval-Augmented Generation. arXiv preprint arXiv:2311.08377, 2023. URL https://arxiv.org/abs/2311.08377. Fangyuan Xu, Weijia Shi, and Eunsol Choi. RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation. In International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=mlJLVigNHp. Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2369–2380, 2018. doi: 10.18653/v1/D18-1259. URL https://aclantholo gy.org/D18-1259/. Zexuan Zhong, Zhengxuan Wu, Christopher Manning, Christopher Potts, and Danqi Chen. MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 15686–15702, 2023. doi: 10.18653/v1/2023.emnlp-main.971. URL https://aclanthology.org/2023. emnlp-main.971/. 12
Text of page 13
A The result tables describe all 150 machine-validated candidates, including the constructiondevelopment, selection, and report partitions. The initial baseline configuration preceded candidate scoring; no settings were tuned on these partitions. Later hosted additions followed inspection of the first results and are identified separately. The earlier construction-control diagnostics remain distinct from this baseline comparison. S TATISTICAL PROCEDURE The supplied uncertainty procedure resamples public source identifiers with replacement, retaining all items associated with each sampled source and the multiplicity of repeated draws. Each replicate computes the item-weighted macro statistic; 10,000 replicates supply the displayed 95% percentile intervals (Cameron et al., 2008). The numeric seed is retained in the release configuration under its designated label. These intervals quantify source-sampling variation under the retained runs; the tables contain no repeat-provider-run uncertainty. The accompanying manifests bind the result exports to the scoring configuration. The item-level records and executable scorer are absent from this handoff, so this manuscript reports the supplied intervals without independently rerunning them. The construction checks and candidate selection used hidden construction labels before model evaluation. The final all-candidate comparison is descriptive; the partition names record construction assignments. B C ONSTRUCTION SCHEMA AND VALIDATION B.1 E DIT , SOURCE UNIT , AND HIDDEN GRAPH The edit record specifies the source unit, an exact old substring, a replacement substring, and a visible description. The edit record itself is separate from the candidate units. The edited source sentence belongs to the constructed dependency closure. Its hidden edge originates at the edit record, so its recorded depth is one. Other affected sentences receive depths from their shortest paths from that record. This convention distinguishes the directly edited source sentence from downstream consequences. The public representation contains document identifiers, titles, unit identifiers, canonical text, and text hashes. Hidden construction metadata stores source provenance, graph edges, template membership, original supporting-fact labels, closure membership, and numeric and entity spans. Systems receive the public representation and edit description. The private fields determine construction checks and later scoring. B.2 S ELECTION AND MECHANICAL CHECKS Source assignment precedes generation and assigns an ordered reserve list to each structural coordinate. The selector takes the first mechanically valid completion in that predetermined list. The aggregate records 1,800 attempted sources, 1,564 mechanical passes, and 150 selected candidates, with 12 reserves per coordinate. Each source eligible for inference receives one completion; failed preflight checks reject a source before generation. The selection rule uses neither baseline scores nor human verdicts. The selected candidates follow the structural schedule rather than a random sample of all mechanical passes. Validation checks the planned slots, graph edges, documents, candidate counts, source grounding, and exact edit substring. It also checks reachability, unaffected units, cross-document consequences, original-support coverage, and canonical text and span hashes. Affected downstream text is composed from a named entity followed by its predicate, with the entity anchored at the start. Unaffected text is generated without that prefix requirement. Internal roles are checked before opaque public identifiers are assigned. Document order and unit order are shuffled independently. The released closure-review protocol specifies how to assess textual support for edges, unintended dependencies, and affected-set completeness. 13
Text of page 14
B.3 Table 4 gives the selected candidate counts for the constructed graph families. Each family contains a branch, and some contain a merge. A merge requires the stated target to follow from each specified parent under the rendered facts. Table 5 summarizes the corresponding document counts, impact-set sizes, and depths. Table 4: Dependency-shape counts on the machine-validated candidate set. Shape 17 23 19 19 18 19 19 16 Table 5: Candidate composition by document count, impact size, and maximum depth on the machine-validated candidate set. Property Value Candidates 3 4 4 5 6 7 3 4 5 75 75 23 56 38 33 42 72 36 Documents Documents Affected units Affected units Affected units Affected units Maximum depth Maximum depth Maximum depth Candidates Branched merge Chain with branch Chain with late branch Deep branch Diamond with tail Unequal-depth branches Fork with tails Three branches D EPENDENCY SHAPES The schedule represents every listed family. Its most frequent shape accounts for the displayed 15.3% of candidates, and its most frequent affected-position pattern accounts for 1.3%. The composition check tests representation, concentration, coverage, and identifier uniqueness. It does not produce a ranked list or a recovery score. C S HORTCUT CONTROLS C.1 C ONTROL DEFINITIONS The identifier control sorts public identifiers in ascending and descending order. For each metric, its diagnostic reports the larger score across these orderings. The document-position control learns affected-unit frequencies from public structural features on development candidates. Its fixed predictor is evaluated separately on selection and report candidates. There is no pooled document-position score in the supplied report. The lexical controls rank by edit-token overlap or BM25 similarity to the edit (Robertson & Zaragoza, 2009). The reported lexical maxima are computed separately for complete recovery and recall. The same row can therefore combine complete recovery from BM25 with recall from token overlap. These component maxima summarize a control family rather than an individual ranking method. The supporting-fact control receives the hidden original-support labels as privileged information. It ranks those units first and orders the remaining units randomly using the fixed seed. It is a 14
Text of page 15
diagnostic of reliance on the source question’s support, rather than a deployable retrieval system. The composition check examines graph-family and position frequencies separately from these ranking controls. C.2 Table 6 pairs complete recovery with recall for every displayed control evaluation. The pooled rows refer to all 150 machine-validated candidates. The random row is the earlier seeded construction diagnostic, whose ordering differs from the retained random baseline in Appendix G. Table 6: Construction-control results on the machine-validated candidate set. Component maxima can come from different rankings. Control Split Identifier max. Identifier max. Identifier max. Identifier max. Document position Document position Lexical maxima Lexical maxima Lexical maxima Lexical maxima Supporting facts Supporting facts Supporting facts Supporting facts Random ordering all 150 development (30) selection (30) report (90) selection (30) report (90) all 150 development (30) selection (30) report (90) all 150 development (30) selection (30) report (90) all 150 Complete (%) Recall@2G 3.3% 10.0% 0.0% 5.6% 3.3% 3.3% 20.0% 13.3% 26.7% 20.0% 4.7% 6.7% 6.7% 3.3% 4.0% 0.56 0.57 0.56 0.57 0.60 0.63 0.68 0.68 0.64 0.69 0.63 0.58 0.63 0.64 0.57 Lexical matching recovers some complete constructed closures, and its selection-split maximum exceeds its pooled value. The registered thresholds are decision rules for detecting the specified construction shortcuts. Passing them establishes the stated checks; semantic acceptance requires the separate review procedure. Table 7 records the ranking thresholds applied to every required evaluation. Table 7: Registered ranking thresholds for complete recovery and recall. Each required split evaluation must remain below both thresholds. C ANDIDATE RESULTS AND THRESHOLDS Control Complete (%) Recall@2G < 35.0 < 35.0 < 60.0 < 35.0 < 0.80 < 0.80 < 0.90 < 0.75 Identifier order Document position Lexical matching Supporting facts The supporting-fact checks additionally require affected units outside the source question’s support and across document boundaries. The generated report gives the non-supporting share of affected units as 81.9%, against the registered minimum of 60.0%. Every item must contain at least two affected units outside the edited document. C.3 W ITHDRAWN CONSTRUCTION An earlier construction gave all 44 development items the same 14-unit structure and four ascending affected identifiers. Identifier order recovered every affected unit without using its text. Those items were withdrawn before audit distribution, and the replacement construction assigned opaque 15
Text of page 16
identifiers after content validation. Later development revisions addressed lexical leakage, coverage
failures, and structured emission, while retaining the registered shortcut thresholds. The surviving
candidate diagnostics evaluate that revised construction.
D
The protocol specifies a review of the constructed answer key against the visible documents. It
separates independent discovery of affected units from inspection of the proposed dependencies.
During blind recovery, two independent reviewers receive identical public documents and the visible
edit. Each selects affected units, gives rationales, checks omissions, and records ambiguity; the
proposed closure remains hidden. Closure review then reveals the proposed membership and graph
edges for independent unit-by-unit and edge-by-edge assessment. Reviewers check textual support,
missing affected statements, unintended dependencies, and ambiguity. Disagreements and unresolved
defects enter joint adjudication after both reviews are committed.
Table 8: Released closure-review procedure. The stages assess target membership and dependency
support against the documents.
R ELEASED CLOSURE - REVIEW PROTOCOL
Step
Evidence required
Blind recovery
Closure review
Adjudication
Independent selections, rationales, and omission checks.
Textual support for membership and dependency edges.
Resolution of membership, edge, and ambiguity questions.
The acceptance rule requires a valid final closure, canonical annotations, source isolation, and no
unresolved membership or dependency question. The procedure preserves independent returns beside
the adjudicated outcome and requires attributable effort records.
E
M ETRIC DEFINITIONS
E.1
R ANKED SELECTION
Let U i be the candidate units and G i ⊆ U i the target affected units for item i. The review cutoff is
K i = min(|U i |, 2|G i |). Let R i (k) contain distinct valid affected identifiers whose first submitted
occurrence lies within the first k positions. Recall at the review cutoff is |R i (K i )|/|G i |. Completeclosure@2G averages the indicator 1{R i (K i ) = G i } over items. Recall@G uses the smaller cutoff
min(|U i |, |G i |).
Unknown identifiers, duplicates, and missing required fields consume positions without receiving
selection credit. A valid affected identifier can receive selection credit when its text is inexact.
Malformed or missing item output is scored as an empty list. Short lists receive no credit for their
missing positions. Identifiers are never inferred from approximate text matches.
Precision at the review cutoff is P i = |R i (K i )|/K i . F-measure combines that precision and recall by
their harmonic mean, with zero assigned when both are zero. Average precision sums the precision at
each first recovered affected identifier and divides by |G i |. Missing affected units contribute zero.
Each item-level statistic is macro-averaged over the evaluation items.
Review burden records the first rank attaining the registered target recall, normalized by impact-set
size. If that target is unreached within |U i | positions, the recorded rank is |U i |+1 before normalization.
The associated success proportion distinguishes unrecovered items from successful review. An allcandidate ranking receives the same prefix evaluation as any other ranking. Its unrestricted recall can
reach complete recovery, while its cutoff scores and burden still depend on ordering.
16
Text of page 17
E.2
Canonical text comparison decodes the submitted JSON string, converts transport-level CRLF line
endings to LF, and compares the resulting UTF-8 bytes. Whitespace, punctuation, Unicode form, and
capitalization otherwise remain unchanged. Verbatim Impact Recall@2G divides correctly copied
affected units within the cutoff by all affected units. Conditional exactness divides correctly copied
recovered units by all recovered affected units, micro-aggregated over the evaluation split. If no
affected identifier is recovered, the conditional score is zero.
Numeric and named-entity annotations identify exact byte intervals in canonical text. A span is
preserved only when its affected identifier is recovered within the cutoff and identical bytes occupy
the annotated interval. Selection-sensitive fidelity divides preserved spans by all annotated affected
spans. Conditional preservation instead divides by annotated spans in recovered affected units.
The result export includes selection-sensitive span fidelity and conditional-preservation fields. The
displayed fidelity table reports the former; conditional verbatim exactness uses recovered affected
units as its denominator. Per-item annotation counts required for full replay are absent from the
supplied aggregate displays. Moved spans, numeric reformatting, aliases, and paraphrases fail this
exact preservation test.
F
T EXT AND SPAN FIDELITY
B ASELINE SPECIFICATIONS AND RETAINED RUNS
The retained initial families comprise six deterministic rankings, BM25 with embedding reranking,
Qwen3 32B, Llama 3.3 70B, and Claude Sonnet 4.5. The reranker uses Titan Text Embeddings V2;
its supplied configuration specifies BM25 retrieval followed by embedding similarity. Every model
receives the same public edit and candidate documents. The system prompt requires all candidate
identifiers and exact source text:
You rank source units affected by a visible edit.
Use only the edit and documents provided. Hidden dependency edges and
, →
gold
membership are unavailable. Rank every candidate unit from most to least
likely affected. Copy each unit's text byte-for-byte from the input.
Return one JSON object and no markdown or commentary:
{"ranking":[{"unit_id":"opaque-id","text":"exact source text"}]}
The requested decoding settings are temperature zero, top-p one, and a 4,096-token output cap. Exact
endpoint identifiers, request records, prompt hashes, and scorer hashes remain in the accompanying
configurations; calendar-bearing identifier suffixes are omitted from this PDF. The initial configuration
specified selection and report evaluation; the revised comparison uses all 150 machine-validated
candidates without tuning on any partition.
Each retained scoring run covers its assigned candidate items once. The campaign also includes
provider-compatibility reruns: the initial Llama result was replaced, and the Sonnet 4.5 result was
replaced twice. The final result manifest identifies which run contributes each row and preserves the
earlier manifests. The pipeline’s no-repair primary scoring rule applies to the retained responses.
F.1
H OSTED FAMILIES ADDED AFTER THE FIRST RESULTS WERE SEEN
Claude Sonnet 5, Claude Opus 5, Claude Fable 5.1, and Grok 4.6 were added under the same prompt,
output cap, parser, and bootstrap procedure. Their provider request schemas reject the specified
temperature field; neither temperature nor top-p is sent, so provider-default sampling applies. Table 9
gives the effective settings, including the thinking-field exception.
17
Text of page 18
Table 9: Provider settings for families added after the first results were seen. All use the same output cap; sampling follows provider defaults. Family Sampling Thinking field Claude Sonnet 5 Claude Opus 5 Claude Fable 5.1 Grok 4.6 Provider default Provider default Provider default Provider default Disabled Disabled Omitted Omitted Fable’s request schema also rejects a disabled-thinking field. G C OMPLETE AND PARTIAL RECOVERY The following scorecards copy the supplied displays at their published precision. Completeclosure@2G and Recall@2G always appear together; Recall@G uses the stricter cutoff. Brackets give the supplied source-cluster intervals. Hosted additions are grouped by their later introduction and provider-default sampling. Table 10: Primary recovery with 95% intervals on the machine-validated candidate set. System Random ordering Render order Identifier ascending Identifier descending Token overlap BM25 BM25 + embeddings Qwen3 32B Llama 3.3 70B Claude Sonnet 4.5 Complete@2G Recall@2G Recall@G 1% [0, 3] 0% [0, 0] 3% [1, 7] 3% [1, 7] 13% [7, 18] 20% [14, 27] 13% [8, 19] 53% [45, 61] 65% [57, 73] 0% [0, 0] 0.54 [0.51, 0.57] 0.57 [0.54, 0.59] 0.55 [0.52, 0.59] 0.56 [0.52, 0.59] 0.68 [0.65, 0.71] 0.66 [0.63, 0.70] 0.68 [0.65, 0.71] 0.72 [0.66, 0.78] 0.91 [0.88, 0.93] 0.00 [0.00, 0.00] 0.27 [0.24, 0.30] 0.28 [0.26, 0.30] 0.27 [0.24, 0.29] 0.29 [0.26, 0.32] 0.39 [0.36, 0.43] 0.37 [0.34, 0.40] 0.43 [0.39, 0.46] 0.64 [0.58, 0.70] 0.78 [0.74, 0.81] 0.00 [0.00, 0.00] Added after the first results were seen; provider-default sampling Claude Sonnet 5 82% [76, 88] 0.95 [0.93, 0.97] 0.89 [0.86, 0.92] Claude Opus 5 0% [0, 0] 0.79 [0.78, 0.81] 0.75 [0.72, 0.78] Claude Fable 5.1 49% [41, 57] 0.89 [0.87, 0.91] 0.85 [0.82, 0.87] Grok 4.6 3% [1, 7] 0.03 [0.01, 0.07] 0.03 [0.01, 0.06] G.1 D EVELOPMENT (30 CANDIDATES ) Table 11: Development (30 candidates) on the machine-validated candidate set. System Random ordering Render order Identifier ascending Identifier descending Token overlap BM25 BM25 + embeddings Qwen3 32B Llama 3.3 70B Claude Sonnet 4.5 Complete@2G Recall@2G Recall@G 0% [0, 0] 0% [0, 0] 10% [0, 20] 0% [0, 0] 7% [0, 17] 13% [3, 27] 10% [0, 23] 57% [40, 73] 70% [53, 87] 0% [0, 0] 0.51 [0.43, 0.58] 0.55 [0.50, 0.60] 0.57 [0.49, 0.65] 0.54 [0.46, 0.62] 0.68 [0.62, 0.74] 0.67 [0.59, 0.74] 0.66 [0.59, 0.73] 0.71 [0.57, 0.85] 0.91 [0.85, 0.96] 0.00 [0.00, 0.00] 0.29 [0.22, 0.35] 0.29 [0.24, 0.33] 0.27 [0.21, 0.32] 0.29 [0.22, 0.36] 0.41 [0.34, 0.47] 0.37 [0.31, 0.44] 0.39 [0.32, 0.47] 0.55 [0.41, 0.69] 0.74 [0.65, 0.83] 0.00 [0.00, 0.00] 18
Text of page 19
Table 11 (continued) System Complete@2G G.2 System Random ordering Render order Identifier ascending Identifier descending Token overlap BM25 BM25 + embeddings Qwen3 32B Llama 3.3 70B Claude Sonnet 4.5 Complete@2G Recall@2G Recall@G 7% [0, 17] 0% [0, 0] 0% [0, 0] 0% [0, 0] 10% [0, 23] 27% [13, 43] 27% [13, 43] 50% [33, 67] 60% [43, 77] 0% [0, 0] 0.59 [0.52, 0.67] 0.56 [0.51, 0.61] 0.56 [0.51, 0.62] 0.53 [0.47, 0.58] 0.64 [0.57, 0.71] 0.64 [0.55, 0.73] 0.77 [0.71, 0.83] 0.65 [0.49, 0.80] 0.91 [0.87, 0.95] 0.00 [0.00, 0.00] 0.30 [0.23, 0.36] 0.27 [0.23, 0.32] 0.30 [0.23, 0.37] 0.28 [0.23, 0.33] 0.36 [0.31, 0.42] 0.33 [0.27, 0.40] 0.51 [0.43, 0.58] 0.62 [0.47, 0.77] 0.79 [0.72, 0.86] 0.00 [0.00, 0.00] Added after the first results were seen; provider-default sampling Claude Sonnet 5 70% [53, 87] 0.93 [0.89, 0.97] 0.89 [0.82, 0.95] Claude Opus 5 0% [0, 0] 0.79 [0.75, 0.81] 0.74 [0.68, 0.79] Claude Fable 5.1 40% [23, 57] 0.86 [0.81, 0.91] 0.83 [0.77, 0.89] Grok 4.6 0% [0, 0] 0.00 [0.00, 0.00] 0.00 [0.00, 0.00] S ELECTION (30 CANDIDATES ) Table 12: Selection (30 candidates) on the machine-validated candidate set. Recall@G Added after the first results were seen; provider-default sampling Claude Sonnet 5 80% [67, 93] 0.95 [0.90, 0.98] 0.87 [0.80, 0.94] Claude Opus 5 0% [0, 0] 0.80 [0.77, 0.82] 0.72 [0.63, 0.80] Claude Fable 5.1 50% [33, 67] 0.89 [0.84, 0.94] 0.83 [0.77, 0.90] Grok 4.6 7% [0, 17] 0.07 [0.00, 0.17] 0.07 [0.00, 0.17] Recall@2G G.3 R EPORT (90 CANDIDATES ) Table 13: Report (90 candidates) on the machine-validated candidate set. System Random ordering Render order Identifier ascending Identifier descending Token overlap BM25 BM25 + embeddings Qwen3 32B Llama 3.3 70B Claude Sonnet 4.5 Complete@2G Recall@2G Recall@G 0% [0, 0] 0% [0, 0] 2% [0, 6] 6% [1, 11] 16% [9, 23] 20% [12, 29] 10% [4, 17] 52% [42, 62] 66% [56, 76] 0% [0, 0] 0.53 [0.49, 0.56] 0.57 [0.54, 0.61] 0.54 [0.50, 0.59] 0.57 [0.53, 0.61] 0.69 [0.65, 0.73] 0.67 [0.62, 0.72] 0.66 [0.61, 0.70] 0.75 [0.67, 0.82] 0.90 [0.87, 0.93] 0.00 [0.00, 0.00] 0.25 [0.21, 0.29] 0.28 [0.26, 0.30] 0.26 [0.22, 0.29] 0.29 [0.26, 0.33] 0.40 [0.35, 0.44] 0.38 [0.34, 0.42] 0.41 [0.37, 0.46] 0.68 [0.60, 0.75] 0.78 [0.73, 0.83] 0.00 [0.00, 0.00] Added after the first results were seen; provider-default sampling Claude Sonnet 5 87% [79, 93] 0.96 [0.93, 0.98] 0.89 [0.85, 0.93] Claude Opus 5 0% [0, 0] 0.79 [0.77, 0.81] 0.76 [0.73, 0.79] Claude Fable 5.1 51% [41, 61] 0.90 [0.87, 0.92] 0.85 [0.82, 0.89] Grok 4.6 3% [0, 8] 0.03 [0.00, 0.08] 0.03 [0.00, 0.07] 19
Text of page 20
H The secondary parser strips one leading code fence and one trailing code fence from the same saved response. It preserves the enclosed content and makes no new model call. The primary parser remains the no-repair evaluation. The saved transformation changes Qwen’s 30 fenced responses and all Sonnet 4.5 responses; the other displayed families have no fence removal. Table 14 retains each primary row beside its secondary counterpart. Table 14: Format-tolerant parse (secondary), alongside primary results, on the machine-validated candidate set. Parse Recall@2G Recall@G Qwen3 32B Primary Fence stripped 120 150 53% [45, 61] 0.72 [0.66, 0.78] 0.64 [0.58, 0.70] 66% [58, 73] 0.90 [0.87, 0.93] 0.80 [0.76, 0.83] Llama 3.3 70B Primary Fence stripped 150 150 65% [57, 73] 0.91 [0.88, 0.93] 0.78 [0.74, 0.81] 65% [57, 73] 0.91 [0.88, 0.93] 0.78 [0.74, 0.81] Claude Sonnet 4.5 Primary 0 Fence stripped 150 0% [0, 0] 0.00 [0.00, 0.00] 0.00 [0.00, 0.00] 33% [26, 41] 0.85 [0.83, 0.87] 0.79 [0.76, 0.82] Claude Sonnet 5 Primary 150 Fence stripped 150 Claude Opus 5 Primary Fence stripped Parsed Complete@2G Added after the first results were seen; provider-default sampling F ORMAT - TOLERANT PARSE ( SECONDARY ) 150 150 0% [0, 0] 0.79 [0.78, 0.81] 0.75 [0.72, 0.78] 0% [0, 0] 0.79 [0.78, 0.81] 0.75 [0.72, 0.78] Claude Fable 5.1 Primary 150 Fence stripped 150 49% [41, 57] 0.89 [0.87, 0.91] 0.85 [0.82, 0.87] 49% [41, 57] 0.89 [0.87, 0.91] 0.85 [0.82, 0.87] Grok 4.6 Primary Fence stripped I 82% [76, 88] 0.95 [0.93, 0.97] 0.89 [0.86, 0.92] 82% [76, 88] 0.95 [0.93, 0.97] 0.89 [0.86, 0.92] 5 5 3% [1, 7] 0.03 [0.01, 0.07] 0.03 [0.01, 0.06] 3% [1, 7] 0.03 [0.01, 0.07] 0.03 [0.01, 0.06] C OPYING AND ANNOTATED - SPAN PRESERVATION Conditional exactness is micro-aggregated over recovered affected units; a system recovering none receives zero. Numeric and named-entity fidelity divide preserved spans by all annotated affected spans, so omissions reduce their scores. These columns therefore measure different denominators rather than alternative estimates of one accuracy. Table 15: Copying and span preservation on the machine-validated candidate set. System Random ordering Render order Identifier ascending Identifier descending Token overlap BM25 BM25 + embeddings Exact copying Numeric spans Entity spans 1.000 1.000 1.000 1.000 1.000 1.000 1.000 20 0.56 0.58 0.54 0.59 0.85 0.91 0.77 0.55 0.57 0.55 0.56 0.68 0.67 0.68
Text of page 21
Table 15 (continued) System Exact copying Numeric spans Entity spans Qwen3 32B Llama 3.3 70B Claude Sonnet 4.5 0.995 0.942 0.000 0.75 0.97 0.00 0.72 0.91 0.00 Added after the first results were seen; provider-default sampling Claude Sonnet 5 0.998 0.94 0.95 Claude Opus 5 0.999 0.28 0.79 Claude Fable 5.1 0.999 0.68 0.89 Grok 4.6 1.000 0.02 0.03 J D IRECT AND DOWNSTREAM RECOVERY The supplied depth export separates the edited source unit, one edge from the visible edit record, from the deeper affected units. Each method has 150 eligible items for both summaries. The scores concern recovery within the same review prefix as the full-set measurements. The export records higher direct than downstream recovery for BM25 and Llama, and the reverse for the later-added Opus and Fable. Opus’s zero direct-recovery score concerns only positions within the reviewed prefix. Empty or malformed responses contribute zero in both summaries. The machinereadable release retains every exported mean and eligible-item count alongside the complete-closure scorecards. The broader registered difficulty slices include document boundaries, original supporting facts, impact size, candidate count, lexical overlap, and annotated spans. No item-level error census or manual taxonomy counts are supplied for those categories. The present interpretation uses the exported depth summaries, parser outcomes, and copying measurements. K S OURCE ISOLATION , LICENSING , AND REPRODUCIBILITY The source is the public MuSiQue release, whose questions combine evidence across Wikipediaderived passages (Trivedi et al., 2022). The supplied record identifies a pinned source revision and archive/member hashes. Source identifiers used by another in-review project were excluded before split assignment, as recorded in the supplementary source-isolation disclosure. The separation record reports disjoint construction-development, selection, and report identifiers, with no transfer of derived translations, model outputs, annotations, or empirical claims. Source-identifier separation defines item provenance; shared passages and prior model exposure remain unassessed. MuSiQue’s Creative Commons Attribution license requires attribution and identification of modifications in source-derived redistribution. New benchmark fields remain distinguishable from upstream material. Available materials contain aggregate results, generated displays, design specifications, audit protocols, request records, and hashes of the referenced inputs. Item-level baseline responses, full private source mappings, executable construction and scoring code, and the original source-isolation records are needed for independent replay. The supplied aggregates permit inspection and display reproduction, while exact result recomputation requires those additional artifacts. L C ONSTRUCTION DECISIONS AFFECTING INTERPRETATION Recorded pilots failed lexical and role/template checks, and one graph family lacked passing candidates. Revisions introduced edit-token restrictions, coverage changes, and structured entity emission. The construction-control report records access to selection and report labels before baseline outputs were available. Baseline evaluation used all candidates without tuning on their partitions. Fenceremoval rescoring and the additional hosted families followed inspection of the first baseline results. Appendix C documents the withdrawn construction and final controls; Appendix F records retained runs and effective settings. The complete construction record remains in the reproducibility archive. 21