telegrapher

LLM labelers score higher when a sibling wrote the answer key: 3.7–3.9 points for Haiku

The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling

Mikhail L Arbuzov, Sisong Bei, Dmitry Dimov, Evgeniya Dontsova, Yaodong Hu, Karan Dave, Vincent Lao, Navita Jain

ICLR 2027 submission, September 2026

An LLM labeler agrees more with silver labels written by a sibling model than with labels from other families. The contrast needs no human gold, so same-family consensus can be audited, and discounted, where gold is scarce.

What we did and found

The setting is an enterprise pipeline that sorts inbound customer-service calls into hundreds of fine-grained call reasons; roughly 76,000 LLM-labeled interactions train a classifier that then runs over tens of millions of conversations. Checking those silver labels against the 924-call human holdout runs into a broken ruler. Every labeler agrees more with LLM silver than with gold, and a manual review suggests part of that gap is genuine gold error, so the gap cannot tell shared model bias apart from models simply being right. The audit keeps gold out of the bias claim instead. It fixes one labeler's predictions, scores them against answer keys written by different models, and reads the difference between a same-family key and a cross-family key. Two leave-out rules keep the contrast honest: a labeler is never scored against its own exact output, and when the key is a panel vote, the labeler is removed from that vote. What remains on the same-family side is a sibling comparison, such as Haiku graded against Sonnet's labels.

Graded against a sibling's key, a labeler scores higher than against a cross-family one. Haiku-4.5 agrees with Sonnet-4.6's labels on 0.838 of calls and with the Gemini, GPT-5.5 and Grok keys on 0.799 to 0.801, a gap of 3.7–3.9 points that survives Bonferroni correction. The pattern holds beyond Haiku: of 20 exploratory (silver key, prompt) contrasts, 17 are significant under Benjamini–Hochberg control and 16 of them are positive, reaching +5.4 points, while the one significant negative comes from a single Grok prompt. A separate statistic agrees, since the Sonnet–Haiku pair fails the same calls more often than Sonnet does with models from other families. Family diversity buys more independence than prompt diversity. Neither buys much; all 48 labelers together behave like fewer than two independent ones. Prompts still matter for accuracy, and the prompt-induced swing widens from the strongest model to the weakest. Diversifying the silver source, including drawing it at random per call, lowers between-family bias dispersion in every construction tried, but at 924 calls every reduction's bootstrap interval includes zero.

Key numbers

Same-family halo, Haiku anchorHaiku-4.5's agreement with Sonnet-4.6's silver labels minus its agreement with each of three cross-family keys; 924 calls, significant after Bonferroni correction3.7–3.9 pp
Exploratory contrasts significant(silver key, prompt) contrasts under Benjamini–Hochberg at q = 0.05; 16 positive, up to +5.4 points17 of 20
Same-family error correlationmean pairwise error correlation of the Sonnet–Haiku pair, against 0.532 for Sonnet's average cross-family pair0.616
Effective independent labelers, full gridM_eff for all 48 labelers together; five prompts on one model give 1.19–1.41, one prompt across six families 1.44–1.581.67
Prompt-induced accuracy swingrange of gold accuracy across five prompts, from the strongest model to the weakest; mean accuracy and spread correlate at r = −0.973.2 to 18.0 points

What this does not show

The evidence comes from one task and one corpus, so the halo's size on other label spaces is untested; the paper conjectures, without testing, that it may appear in medical coding, moderation taxonomies and LLM juries. Within the panel, two vendor lines contribute sibling models (Anthropic and OpenAI), and the effect is most directly evidenced for the Anthropic pair. Same-family and cross-family sources are not matched on capability, so capability proximity rather than lineage may drive part of the halo; the paper flags this confound and does not resolve it. The 20 exploratory contrasts share the same calls and overlapping models and prompts, so they corroborate the headline rather than replicate it. The effective-labeler count is an approximation, not an exact count. The randomized construction's benefit is directional and measured against gold, with every bootstrap interval including zero at 924 calls. Nor does the paper show that a family-diverse panel yields a more accurate downstream classifier, or settle how much of the general silver–gold divergence is shared bias and how much is models being right where humans were wrong. The call data cannot be released.

Every number above was checked against the paper text.
The authors' abstract

Scalable analysis of long-form human–human, human–agent, and agent–agent interactions requires reliable supervision. Large language models (LLMs) provide a practical source of silver labels, but agreement among models can reflect shared labeling preferences rather than independent confirmation. We investigate this dependence in an enterprise pipeline that uses 76,000 silver-labeled customerservice interactions to train classifiers operating over tens of millions of conversations, with a separate human-labeled holdout of 924 examples. We introduce a gold-free audit: holding each labeler’s predictions fixed, we vary the model supplying the reference labels and measure changes in agreement without consulting human annotations. Across 48 model–prompt–rendering configurations, we observe source-dependent agreement that extends beyond exact self-comparisons to sibling models. In a representative comparison, agreement with sibling-model labels exceeds agreement with three cross-family sources by 3.7–3.9 percentage points. To mitigate this effect, we propose combinatorial silver-label construction, assigning each example to a randomly selected model–prompt tuple — randomizing the prompt as well, since prompt choice is itself a first-order driver of label quality. Under this construction, the source-specific agreement advantage seen with fixed-source labels is directionally reduced in our evaluation, though not to statistical significance at our sample size. The method distributes supervision across configurations while preserving exactly one inference call per example. Our central — and demonstrated — finding is that same-family consensus can reflect a source-dependent agreement pattern that resembles independent confirmation while not providing it. The gold-free audit makes this dependence measurable even where human reference labels are scarce, and randomized construction offers a practical, single-call route toward mitigating it in scalable conversation analytics.

Figures

Figure 1 of the paper.
Figure 1 of the paper.
Figure 2 of the paper.
Figure 2 of the paper.
On the blogAn LLM labeler scores higher when its sibling wrote the answer keyHold an LLM labeler fixed and swap who wrote its answer key: agreement rises when the key comes from a sibling model, and no human gold is needed to see it.

More in Conversation analytics

Cite

@misc{arbuzov2026the,
  title         = {The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Dave, Karan and Lao, Vincent and Jain, Navita},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/same-family-halo/}
}

Builds on

  1. Kim et al. (2025). Correlated errors in large language models.
  2. Panickssery et al. (2024). LLM evaluators recognize and favor their own generations.
  3. Wataoka et al. (2024). Self-preference bias in LLM-as-a-judge.
  4. Dawid and Skene (1979). Maximum likelihood estimation of observer error-rates using the EM algorithm.
  5. Krogh and Vedelsby (1994). Neural network ensembles, cross validation, and active learning.