{
  "slug": "same-family-halo",
  "title": "The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling",
  "short": "Halo",
  "line": "applied",
  "line_name": "Conversation analytics",
  "part": null,
  "status": "ICLR 2027 submission",
  "date": "2026-09-17",
  "authors": [
    "Mikhail L Arbuzov",
    "Sisong Bei",
    "Dmitry Dimov",
    "Evgeniya Dontsova",
    "Yaodong Hu",
    "Karan Dave",
    "Vincent Lao",
    "Navita Jain"
  ],
  "abstract": "Scalable analysis of long-form human–human, human–agent, and agent–agent interactions requires reliable supervision. Large language models (LLMs) provide a practical source of silver labels, but agreement among models can reflect shared labeling preferences rather than independent confirmation. We investigate this dependence in an enterprise pipeline that uses 76,000 silver-labeled customerservice interactions to train classifiers operating over tens of millions of conversations, with a separate human-labeled holdout of 924 examples. We introduce a gold-free audit: holding each labeler’s predictions fixed, we vary the model supplying the reference labels and measure changes in agreement without consulting human annotations. Across 48 model–prompt–rendering configurations, we observe source-dependent agreement that extends beyond exact self-comparisons to sibling models. In a representative comparison, agreement with sibling-model labels exceeds agreement with three cross-family sources by 3.7–3.9 percentage points. To mitigate this effect, we propose combinatorial silver-label construction, assigning each example to a randomly selected model–prompt tuple — randomizing the prompt as well, since prompt choice is itself a first-order driver of label quality. Under this construction, the source-specific agreement advantage seen with fixed-source labels is directionally reduced in our evaluation, though not to statistical significance at our sample size. The method distributes supervision across configurations while preserving exactly one inference call per example. Our central — and demonstrated — finding is that same-family consensus can reflect a source-dependent agreement pattern that resembles independent confirmation while not providing it. The gold-free audit makes this dependence measurable even where human reference labels are scarce, and randomized construction offers a practical, single-call route toward mitigating it in scalable conversation analytics.",
  "tldr": "An LLM labeler agrees more with silver labels written by a sibling model than with labels from other families. The contrast needs no human gold, so same-family consensus can be audited, and discounted, where gold is scarce.",
  "pages": 17,
  "html": "https://telegrapher.ai/research/same-family-halo/",
  "md": "https://telegrapher.ai/research/same-family-halo.md",
  "reader": "https://telegrapher.ai/research/same-family-halo/read/",
  "pdf": "https://telegrapher.ai/papers/same-family-halo/same-family-halo.pdf",
  "arxiv": null,
  "openreview": "https://openreview.net/forum?id=7a5Il6NjMv",
  "post": "https://telegrapher.ai/blog/same-family-halo/",
  "bibtex": "@misc{arbuzov2026the,\n  title         = {The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling},\n  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Dave, Karan and Lao, Vincent and Jain, Navita},\n  year          = {2026},\n  note          = {ICLR 2027 submission},\n  url           = {https://telegrapher.ai/research/same-family-halo/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "LLM labelers score higher when a sibling wrote the answer key: 3.7–3.9 points for Haiku"
    },
    {
      "label": "TL;DR",
      "text": "An LLM labeler agrees more with silver labels written by a sibling model than with labels from other families. The contrast needs no human gold, so same-family consensus can be audited, and discounted, where gold is scarce."
    },
    {
      "label": "Method",
      "text": "A grid of 48 labeler configurations (model, prompt, and whether category definitions are shown) labels the same 924 human-gold customer-service calls; each labeler is scored against silver answer keys written by same-family and cross-family models, the within-labeler difference is tested with paired McNemar tests, and panels are compared by mean pairwise error correlation and an effective-labeler count, M_eff."
    },
    {
      "label": "Key result",
      "text": "Same-family halo, Haiku anchor: 3.7–3.9 pp; Exploratory contrasts significant: 17 of 20; Same-family error correlation: 0.616"
    },
    {
      "label": "Why it matters",
      "text": "This matters wherever LLM agreement is used to certify LLM labels."
    },
    {
      "label": "Limits",
      "text": "The evidence comes from one task and one corpus, so the halo's size on other label spaces is untested; the paper conjectures, without testing, that it may appear in medical coding, moderation taxonomies and LLM juries."
    },
    {
      "label": "Status",
      "text": "ICLR 2027 submission, September 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/same-family-halo/read/, PDF /papers/same-family-halo/same-family-halo.pdf"
    }
  ],
  "note": {
    "slug": "same-family-halo",
    "claim_title": "LLM labelers score higher when a sibling wrote the answer key: 3.7–3.9 points for Haiku",
    "meta_description": "Hold an LLM labeler fixed, swap the model that wrote its answer key, and agreement rises for a sibling. The audit needs no human gold labels.",
    "tldr": "An LLM labeler agrees more with silver labels written by a sibling model than with labels from other families. The contrast needs no human gold, so same-family consensus can be audited, and discounted, where gold is scarce.",
    "gist": "Pipelines that train classifiers on LLM-written silver labels often vet those labels by agreement: when several models concur, the label is trusted. That reading assumes the models err independently. Human gold cannot settle the question, because the gold is itself sometimes wrong. So the paper takes gold out of the bias claim. It holds one labeler's predictions fixed and changes only the model that wrote the answer key; the labeler's own competence is the same on both sides and cancels. On 924 customer-service calls labeled into hundreds of fine-grained call reasons, Haiku agrees with Sonnet's labels 3.7–3.9 percentage points more than with labels from Gemini, GPT-5.5 or Grok. Across 20 exploratory contrasts, 17 are significant and 16 of those are positive. Sibling models also fail the same calls more often, and spreading a panel across model families decorrelates its errors more than spreading it across prompts. The proposed fix draws each call's silver source at random, at one inference call per label. Its benefit is directional, not statistically significant at this sample size.",
    "method": "A grid of 48 labeler configurations (model, prompt, and whether category definitions are shown) labels the same 924 human-gold customer-service calls; each labeler is scored against silver answer keys written by same-family and cross-family models, the within-labeler difference is tested with paired McNemar tests, and panels are compared by mean pairwise error correlation and an effective-labeler count, M_eff.",
    "summary_html": [
      "The setting is an enterprise pipeline that sorts inbound customer-service calls into hundreds of fine-grained call reasons; roughly 76,000 LLM-labeled interactions train a classifier that then runs over tens of millions of conversations. Checking those silver labels against the 924-call human holdout runs into a <em>broken ruler</em>. Every labeler agrees more with LLM silver than with gold, and a manual review suggests part of that gap is genuine gold error, so the gap cannot tell shared model bias apart from models simply being right. The audit keeps gold out of the bias claim instead. It fixes one labeler's predictions, scores them against answer keys written by different models, and reads the difference between a same-family key and a cross-family key. Two leave-out rules keep the contrast honest: a labeler is never scored against its own exact output, and when the key is a panel vote, the labeler is removed from that vote. What remains on the same-family side is a sibling comparison, such as Haiku graded against Sonnet's labels.",
      "Graded against a sibling's key, a labeler scores higher than against a cross-family one. Haiku-4.5 agrees with Sonnet-4.6's labels on 0.838 of calls and with the Gemini, GPT-5.5 and Grok keys on 0.799 to 0.801, a gap of 3.7–3.9 points that survives Bonferroni correction. The pattern holds beyond Haiku: of 20 exploratory (silver key, prompt) contrasts, 17 are significant under Benjamini–Hochberg control and 16 of them are positive, reaching +5.4 points, while the one significant negative comes from a single Grok prompt. A separate statistic agrees, since the Sonnet–Haiku pair fails the same calls more often than Sonnet does with models from other families. Family diversity buys more independence than prompt diversity. Neither buys much; all 48 labelers together behave like fewer than two independent ones. Prompts still matter for accuracy, and the prompt-induced swing widens from the strongest model to the weakest. Diversifying the silver source, including drawing it at random per call, lowers between-family bias dispersion in every construction tried, but at 924 calls every reduction's bootstrap interval includes zero."
    ],
    "key_numbers": [
      {
        "label": "Same-family halo, Haiku anchor",
        "value": "3.7–3.9 pp",
        "context": "Haiku-4.5's agreement with Sonnet-4.6's silver labels minus its agreement with each of three cross-family keys; 924 calls, significant after Bonferroni correction",
        "bad": true
      },
      {
        "label": "Exploratory contrasts significant",
        "value": "17 of 20",
        "context": "(silver key, prompt) contrasts under Benjamini–Hochberg at q = 0.05; 16 positive, up to +5.4 points"
      },
      {
        "label": "Same-family error correlation",
        "value": "0.616",
        "context": "mean pairwise error correlation of the Sonnet–Haiku pair, against 0.532 for Sonnet's average cross-family pair",
        "bad": true
      },
      {
        "label": "Effective independent labelers, full grid",
        "value": "1.67",
        "context": "M_eff for all 48 labelers together; five prompts on one model give 1.19–1.41, one prompt across six families 1.44–1.58",
        "bad": true
      },
      {
        "label": "Prompt-induced accuracy swing",
        "value": "3.2 to 18.0 points",
        "context": "range of gold accuracy across five prompts, from the strongest model to the weakest; mean accuracy and spread correlate at r = −0.97"
      }
    ],
    "editorial_html": [
      "This matters wherever LLM agreement is used to certify LLM labels. Silver-label training sets are vetted that way, and the label models of weak supervision classically assume that sources err independently once the true label is known. The halo breaks that assumption in a structured way, along family lines, so a same-family consensus claims more confidence than its evidence supports. The paper's guidance follows directly. A panel's diversity budget is better spent on model families than on prompt variants of one model, which tend to fail the same calls. No single fixed source should write the answer key, because whichever family authors it is the one whose siblings get flattered. And agreement between siblings should not count as stronger confirmation than agreement across families.",
      "None of this makes prompt choice cheap. A prompt barely moves a strong labeler and moves a weak one a great deal, which is why the proposed construction randomizes the prompt along with the model: a corpus built from one prompt inherits that prompt's quality profile. The call count stays flat, one inference call per label, against the several a voting panel spends.",
      "The paper sits in the group's conversation-analytics line, on the same kind of fine-grained call-reason taxonomy. <em>High-Load Budgeted Categorization of Customer Care Calls</em> describes a deployed encoder–LLM cascade that routes customer-call summaries into fine-grained, long-tailed categories; this paper asks whether the labels a classifier of that kind learns from deserve trust. Disagreement among labelers here collapses onto a few pairs of semantically adjacent categories, which points back at the category definitions. <em>Clusters Are Proposals</em> comes at the taxonomy from the other end: it discovers finer subcategories inside broad call groups, gives each candidate a written definition, and merges two candidates when an LLM confirms they describe the same subcategory."
    ],
    "limitations": "The evidence comes from one task and one corpus, so the halo's size on other label spaces is untested; the paper conjectures, without testing, that it may appear in medical coding, moderation taxonomies and LLM juries. Within the panel, two vendor lines contribute sibling models (Anthropic and OpenAI), and the effect is most directly evidenced for the Anthropic pair. Same-family and cross-family sources are not matched on capability, so capability proximity rather than lineage may drive part of the halo; the paper flags this confound and does not resolve it. The 20 exploratory contrasts share the same calls and overlapping models and prompts, so they corroborate the headline rather than replicate it. The effective-labeler count is an approximation, not an exact count. The randomized construction's benefit is directional and measured against gold, with every bootstrap interval including zero at 924 calls. Nor does the paper show that a family-diverse panel yields a more accurate downstream classifier, or settle how much of the general silver–gold divergence is shared bias and how much is models being right where humans were wrong. The call data cannot be released.",
    "concept_terms": [
      "same-family halo",
      "silver labels",
      "gold-free audit",
      "LLM labeler agreement",
      "error decorrelation",
      "effective number of labelers"
    ],
    "references": [
      "Kim et al. (2025). Correlated errors in large language models.",
      "Panickssery et al. (2024). LLM evaluators recognize and favor their own generations.",
      "Wataoka et al. (2024). Self-preference bias in LLM-as-a-judge.",
      "Dawid and Skene (1979). Maximum likelihood estimation of observer error-rates using the EM algorithm.",
      "Krogh and Vedelsby (1994). Neural network ensembles, cross validation, and active learning."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/same-family-halo.txt",
    "verification": {
      "claims": [
        {
          "claim": "claim_title / gist / key_numbers: Haiku's agreement with a sibling's (Sonnet) silver labels exceeds its agreement with three cross-family keys by 3.7–3.9 percentage points (Table 1 H = +0.039, +0.039, +0.037; recomputed 0.838−0.799 = 0.039, 0.838−0.801 = 0.037)",
          "verdict": "CONFIRMED",
          "source_quote": "exceeds agreement with three cross-family sources by 3.7–3.9 percentage"
        },
        {
          "claim": "summary: Haiku-4.5 agrees with Sonnet-4.6's key on 0.838 of calls and with the Gemini-2.5-Pro, GPT-5.5 and Grok-4.6 keys on 0.799 / 0.799 / 0.801 (Table 1 cells)",
          "verdict": "CONFIRMED",
          "source_quote": "Holding Haiku-4.5 as a"
        },
        {
          "claim": "summary: the same-family key is Sonnet-4.6 (Anthropic); the cross-family keys are Google, OpenAI and xAI models",
          "verdict": "CONFIRMED",
          "source_quote": "its agreement with the same-family (Sonnet-4.6) silver key is compared against its"
        },
        {
          "claim": "summary / key_numbers: the Haiku gap survives Bonferroni correction (Table 1 corrected p = 7.9e-4, 7.9e-4, 2.3e-3)",
          "verdict": "CONFIRMED",
          "source_quote": "surviving Bonferroni correction (Table 1)"
        },
        {
          "claim": "claim_title / tldr / summary: holding the labeler fixed, a labeler scores higher against a same-family (sibling) key than against a cross-family key (Finding 1)",
          "verdict": "CONFIRMED",
          "source_quote": "Holding the labeler fixed, a labeler"
        },
        {
          "claim": "summary: the pattern holds beyond Haiku (Figure 2: the halo generalizes across 20 (silver, prompt) contrasts)",
          "verdict": "CONFIRMED",
          "source_quote": "The halo generalizes across 20 (silver, prompt) contrasts"
        },
        {
          "claim": "gist / summary / key_numbers: 17 of 20 exploratory (silver key, prompt) contrasts are significant under Benjamini–Hochberg at q = 0.05",
          "verdict": "CONFIRMED",
          "source_quote": "Under Benjamini–Hochberg control at q = 0.05, 17 of 20"
        },
        {
          "claim": "gist / summary / key_numbers: 16 of the significant contrasts are positive, up to +5.4 points (derived: 17 significant minus the single significant negative = 16)",
          "verdict": "CONFIRMED",
          "source_quote": "16 are positive (the halo, up to +5.4 points). The single"
        },
        {
          "claim": "summary: the one significant negative contrast is a single Grok prompt",
          "verdict": "CONFIRMED",
          "source_quote": "significant negative contrast (Grok / distinctive-evidence"
        },
        {
          "claim": "limitations: the 20 contrasts share calls and overlapping models and prompts, so they corroborate rather than replicate the headline",
          "verdict": "CONFIRMED",
          "source_quote": "so we read them as repeated internal contrasts that corroborate the headline, not as twenty separate"
        },
        {
          "claim": "gist / summary / key_numbers: the Sonnet–Haiku pair has mean pairwise error correlation 0.616 vs 0.532 for Sonnet's average cross-family pair (sibling models fail the same calls more often)",
          "verdict": "CONFIRMED",
          "source_quote": "0.616 (M eff = 1.24), versus"
        },
        {
          "claim": "key_numbers: 0.532 is Sonnet's average cross-family pair",
          "verdict": "CONFIRMED",
          "source_quote": "average Sonnet↔cross-family pair"
        },
        {
          "claim": "summary / key_numbers: all 48 labelers together give M_eff 1.67, behaving like fewer than two independent labelers",
          "verdict": "CONFIRMED",
          "source_quote": "the full grid of 48 labelers still behaves like fewer than two"
        },
        {
          "claim": "key_numbers: five prompts on one model give M_eff 1.19–1.41",
          "verdict": "CONFIRMED",
          "source_quote": "on a single model give M eff ≈ 1.19–1.41"
        },
        {
          "claim": "key_numbers: one prompt across six families gives M_eff 1.44–1.58",
          "verdict": "CONFIRMED",
          "source_quote": "whereas one prompt across six families gives M eff ≈ 1.44–1.58"
        },
        {
          "claim": "gist / summary / editorial: family diversity decorrelates errors more than prompt diversity, but neither buys much",
          "verdict": "CONFIRMED",
          "source_quote": "margin is narrow, and neither buys much in absolute terms"
        },
        {
          "claim": "editorial: prompt variants of one model tend to fail the same calls",
          "verdict": "CONFIRMED",
          "source_quote": "model tend to fail the same calls"
        },
        {
          "claim": "summary / key_numbers: the prompt-induced gold-accuracy swing runs from 3.2 points on the strongest model to 18.0 on the weakest",
          "verdict": "CONFIRMED",
          "source_quote": "3.2 points on the strongest model to 18.0 on the weakest (Table 3)"
        },
        {
          "claim": "key_numbers: mean accuracy and prompt spread correlate at r = −0.97 (recomputed from Table 3: −0.967)",
          "verdict": "CONFIRMED",
          "source_quote": "mean accuracy and prompt-induced spread correlate at r = −0.97"
        },
        {
          "claim": "editorial: a prompt barely moves a strong labeler and moves a weak one a great deal",
          "verdict": "CONFIRMED",
          "source_quote": "It barely moves a strong labeler but is a first-order"
        },
        {
          "claim": "summary: every labeler agrees more with LLM silver than with gold",
          "verdict": "CONFIRMED",
          "source_quote": "against LLM silver than against gold labels"
        },
        {
          "claim": "gist / summary: a manual review suggests part of the silver–gold gap is genuine gold error; gold is itself sometimes wrong",
          "verdict": "CONFIRMED",
          "source_quote": "is genuine gold error — human labels that are flatly wrong where the model is right"
        },
        {
          "claim": "summary: roughly 76,000 LLM-labeled interactions train a classifier that runs over tens of millions of conversations",
          "verdict": "CONFIRMED",
          "source_quote": "roughly 76,000 silver-labeled interactions train a downstream classifier"
        },
        {
          "claim": "summary / gist / method: a 924-call human-labeled holdout",
          "verdict": "CONFIRMED",
          "source_quote": "(N = 924) available to check the silver against"
        },
        {
          "claim": "gist / summary: calls are labeled into hundreds of fine-grained call reasons",
          "verdict": "CONFIRMED",
          "source_quote": "The finest level contains hundreds of fine-grained call-reason categories"
        },
        {
          "claim": "gist / summary: the audit holds one labeler fixed and varies only who wrote the answer key, so the labeler's competence cancels",
          "verdict": "CONFIRMED",
          "source_quote": "a single labeler fixed and vary only whose labels it is scored against"
        },
        {
          "claim": "summary: leave-out rule 1, a labeler is never scored against its own exact output",
          "verdict": "CONFIRMED",
          "source_quote": "a labeler scored against its own exact output, which is 100% by construction"
        },
        {
          "claim": "summary: leave-out rule 2, when the key is a panel vote the labeler is removed from that vote",
          "verdict": "CONFIRMED",
          "source_quote": "when the silver answer key is a panel the labeler is removed from that panel"
        },
        {
          "claim": "summary: the same-family side is a sibling comparison, e.g., Haiku graded against Sonnet's labels",
          "verdict": "CONFIRMED",
          "source_quote": "e.g., Haiku against Sonnet"
        },
        {
          "claim": "method: 48 labeler configurations of model, prompt and rendering on the same 924 calls",
          "verdict": "CONFIRMED",
          "source_quote": "per family and 48 labelers in total"
        },
        {
          "claim": "method: rendering = whether category definitions are shown",
          "verdict": "CONFIRMED",
          "source_quote": "settings (full vs. no category definitions; see Appendix C)"
        },
        {
          "claim": "method: the within-labeler difference is tested with paired McNemar tests",
          "verdict": "CONFIRMED",
          "source_quote": "H is tested with the paired McNemar"
        },
        {
          "claim": "limitations: M_eff is an approximation, not an exact count",
          "verdict": "CONFIRMED",
          "source_quote": "approximation for correlated binary error indicators, not an exact count"
        },
        {
          "claim": "gist / editorial: the proposed fix draws each call's silver source at random, with the prompt randomized alongside the model",
          "verdict": "CONFIRMED",
          "source_quote": "we draw the source at random per call"
        },
        {
          "claim": "gist / editorial: one inference call per label, against K calls for a voting panel",
          "verdict": "CONFIRMED",
          "source_quote": "it uses exactly one model call per label"
        },
        {
          "claim": "editorial: a single fixed prompt binds the corpus to that prompt's quality profile",
          "verdict": "CONFIRMED",
          "source_quote": "so fixing one prompt would bind the whole"
        },
        {
          "claim": "editorial: whichever family authors a fixed answer key is the one whose siblings get flattered",
          "verdict": "CONFIRMED",
          "source_quote": "fixed silver source hands a systematic advantage to its own family"
        },
        {
          "claim": "summary: every diversified construction lowers between-family bias dispersion below a single fixed source (Table 4: all D below 0.0172)",
          "verdict": "CONFIRMED",
          "source_quote": "construction we tested reduces it below a single fixed silver source"
        },
        {
          "claim": "gist / summary / limitations: the construction's benefit is directional; at 924 calls every bootstrap interval includes zero, and the measure is gold-referenced",
          "verdict": "CONFIRMED",
          "source_quote": "calls all ∆D bootstrap intervals include zero"
        },
        {
          "claim": "editorial: siblings' shared correlated errors survive the construction",
          "verdict": "CONFIRMED",
          "source_quote": "correlated errors that two models of a family make together"
        },
        {
          "claim": "editorial: weak-supervision label models classically assume sources err independently given the true label",
          "verdict": "CONFIRMED",
          "source_quote": "Weak supervision and consensus label models estimate latent truth from multiple noisy sources"
        },
        {
          "claim": "editorial: a same-family consensus claims more confidence than its evidence supports",
          "verdict": "CONFIRMED",
          "source_quote": "a same-family consensus is over-confident"
        },
        {
          "claim": "editorial: agreement between siblings should not count as stronger confirmation than agreement across families",
          "verdict": "CONFIRMED",
          "source_quote": "do not treat same-family consensus as stronger independent confirmation than cross-family"
        },
        {
          "claim": "editorial: labeler disagreement collapses onto a few semantically adjacent category pairs",
          "verdict": "CONFIRMED",
          "source_quote": "disagreement concentrates in a few semantically adjacent category pairs"
        },
        {
          "claim": "limitations: one task and one corpus; the halo's size on other label spaces is untested",
          "verdict": "CONFIRMED",
          "source_quote": "a single task and corpus"
        },
        {
          "claim": "limitations: the paper conjectures, without testing, medical coding, moderation taxonomies and LLM juries",
          "verdict": "CONFIRMED",
          "source_quote": "We conjecture, rather than demonstrate, that similar source dependence may arise wherever"
        },
        {
          "claim": "limitations: two vendor lines (Anthropic and OpenAI) contribute sibling models",
          "verdict": "CONFIRMED",
          "source_quote": "models (Anthropic and OpenAI), so the within-panel same-family contrasts are concentrated"
        },
        {
          "claim": "limitations: the effect is most directly evidenced for the Anthropic pair",
          "verdict": "CONFIRMED",
          "source_quote": "the effect is most directly evidenced for the Anthropic pair"
        },
        {
          "claim": "limitations: sources are not matched on capability; lineage vs capability flagged, not resolved",
          "verdict": "CONFIRMED",
          "source_quote": "same- and cross-family silver sources are not matched on capability"
        },
        {
          "claim": "limitations: no demonstration that a family-diverse panel yields a more accurate downstream model",
          "verdict": "CONFIRMED",
          "source_quote": "we do not demonstrate that a family-diverse panel yields a"
        },
        {
          "claim": "limitations: the call data cannot be released",
          "verdict": "CONFIRMED",
          "source_quote": "The data are real customer-service call transcripts and cannot be released"
        },
        {
          "claim": "editorial (other paper, checked against abstracts.yaml): High-Load Budgeted Categorization describes a deployed encoder–LLM cascade routing customer-call summaries into fine-grained, long-tailed categories",
          "verdict": "CONFIRMED",
          "source_quote": "We present a deployed hybrid encoder–LLM cascade."
        },
        {
          "claim": "editorial (other paper, checked against abstracts.yaml): Clusters Are Proposals finds finer subcategories inside broad call groups, each candidate with a written definition, merged when an LLM confirms they describe the same subcategory",
          "verdict": "CONFIRMED",
          "source_quote": "Candidates are merged only when an LLM confirms, from their definitions and supporting texts, that they describe the same subcategory"
        },
        {
          "claim": "references: all five entries (Kim et al. 2025; Panickssery et al. 2024; Wataoka et al. 2024; Dawid and Skene 1979; Krogh and Vedelsby 1994) are in the paper's reference list",
          "verdict": "CONFIRMED",
          "source_quote": "Correlated errors in large language"
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}