---
type: paper
slug: same-family-halo
title: 'The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM
  Silver Labeling'
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Dmitry Dimov
- Evgeniya Dontsova
- Yaodong Hu
- Karan Dave
- Vincent Lao
- Navita Jain
date: '2026-09-17'
status: ICLR 2027 submission
line: Conversation analytics
pages: 17
html: https://telegrapher.ai/research/same-family-halo/
pdf: https://telegrapher.ai/papers/same-family-halo/same-family-halo.pdf
reader: https://telegrapher.ai/research/same-family-halo/read/
json: https://telegrapher.ai/api/papers/same-family-halo.json
openreview: https://openreview.net/forum?id=7a5Il6NjMv
---

# The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling

## Paper gist

- **Claim:** LLM labelers score higher when a sibling wrote the answer key: 3.7–3.9 points for Haiku
- **TL;DR:** An LLM labeler agrees more with silver labels written by a sibling model than with labels from other families. The contrast needs no human gold, so same-family consensus can be audited, and discounted, where gold is scarce.
- **Method:** A grid of 48 labeler configurations (model, prompt, and whether category definitions are shown) labels the same 924 human-gold customer-service calls; each labeler is scored against silver answer keys written by same-family and cross-family models, the within-labeler difference is tested with paired McNemar tests, and panels are compared by mean pairwise error correlation and an effective-labeler count, M_eff.
- **Key result:** Same-family halo, Haiku anchor: 3.7–3.9 pp; Exploratory contrasts significant: 17 of 20; Same-family error correlation: 0.616
- **Why it matters:** This matters wherever LLM agreement is used to certify LLM labels.
- **Limits:** The evidence comes from one task and one corpus, so the halo's size on other label spaces is untested; the paper conjectures, without testing, that it may appear in medical coding, moderation taxonomies and LLM juries.
- **Status:** ICLR 2027 submission, September 2026
- **Read:** reader /research/same-family-halo/read/, PDF /papers/same-family-halo/same-family-halo.pdf

## Abstract

Scalable analysis of long-form human–human, human–agent, and agent–agent interactions requires reliable supervision. Large language models (LLMs) provide a practical source of silver labels, but agreement among models can reflect shared labeling preferences rather than independent confirmation. We investigate this dependence in an enterprise pipeline that uses 76,000 silver-labeled customerservice interactions to train classifiers operating over tens of millions of conversations, with a separate human-labeled holdout of 924 examples. We introduce a gold-free audit: holding each labeler’s predictions fixed, we vary the model supplying the reference labels and measure changes in agreement without consulting human annotations. Across 48 model–prompt–rendering configurations, we observe source-dependent agreement that extends beyond exact self-comparisons to sibling models. In a representative comparison, agreement with sibling-model labels exceeds agreement with three cross-family sources by 3.7–3.9 percentage points. To mitigate this effect, we propose combinatorial silver-label construction, assigning each example to a randomly selected model–prompt tuple — randomizing the prompt as well, since prompt choice is itself a first-order driver of label quality. Under this construction, the source-specific agreement advantage seen with fixed-source labels is directionally reduced in our evaluation, though not to statistical significance at our sample size. The method distributes supervision across configurations while preserving exactly one inference call per example. Our central — and demonstrated — finding is that same-family consensus can reflect a source-dependent agreement pattern that resembles independent confirmation while not providing it. The gold-free audit makes this dependence measurable even where human reference labels are scarce, and randomized construction offers a practical, single-call route toward mitigating it in scalable conversation analytics.

## What we did and found

The setting is an enterprise pipeline that sorts inbound customer-service calls into hundreds of fine-grained call reasons; roughly 76,000 LLM-labeled interactions train a classifier that then runs over tens of millions of conversations. Checking those silver labels against the 924-call human holdout runs into a broken ruler. Every labeler agrees more with LLM silver than with gold, and a manual review suggests part of that gap is genuine gold error, so the gap cannot tell shared model bias apart from models simply being right. The audit keeps gold out of the bias claim instead. It fixes one labeler's predictions, scores them against answer keys written by different models, and reads the difference between a same-family key and a cross-family key. Two leave-out rules keep the contrast honest: a labeler is never scored against its own exact output, and when the key is a panel vote, the labeler is removed from that vote. What remains on the same-family side is a sibling comparison, such as Haiku graded against Sonnet's labels.

Graded against a sibling's key, a labeler scores higher than against a cross-family one. Haiku-4.5 agrees with Sonnet-4.6's labels on 0.838 of calls and with the Gemini, GPT-5.5 and Grok keys on 0.799 to 0.801, a gap of 3.7–3.9 points that survives Bonferroni correction. The pattern holds beyond Haiku: of 20 exploratory (silver key, prompt) contrasts, 17 are significant under Benjamini–Hochberg control and 16 of them are positive, reaching +5.4 points, while the one significant negative comes from a single Grok prompt. A separate statistic agrees, since the Sonnet–Haiku pair fails the same calls more often than Sonnet does with models from other families. Family diversity buys more independence than prompt diversity. Neither buys much; all 48 labelers together behave like fewer than two independent ones. Prompts still matter for accuracy, and the prompt-induced swing widens from the strongest model to the weakest. Diversifying the silver source, including drawing it at random per call, lowers between-family bias dispersion in every construction tried, but at 924 calls every reduction's bootstrap interval includes zero.

## Key numbers

| Measure | Value |
|---|---|
| Same-family halo, Haiku anchor | 3.7–3.9 pp |
| Exploratory contrasts significant | 17 of 20 |
| Same-family error correlation | 0.616 |
| Effective independent labelers, full grid | 1.67 |
| Prompt-induced accuracy swing | 3.2 to 18.0 points |

## Why it matters

This matters wherever LLM agreement is used to certify LLM labels. Silver-label training sets are vetted that way, and the label models of weak supervision classically assume that sources err independently once the true label is known. The halo breaks that assumption in a structured way, along family lines, so a same-family consensus claims more confidence than its evidence supports. The paper's guidance follows directly. A panel's diversity budget is better spent on model families than on prompt variants of one model, which tend to fail the same calls. No single fixed source should write the answer key, because whichever family authors it is the one whose siblings get flattered. And agreement between siblings should not count as stronger confirmation than agreement across families.

None of this makes prompt choice cheap. A prompt barely moves a strong labeler and moves a weak one a great deal, which is why the proposed construction randomizes the prompt along with the model: a corpus built from one prompt inherits that prompt's quality profile. The call count stays flat, one inference call per label, against the several a voting panel spends.

The paper sits in the group's conversation-analytics line, on the same kind of fine-grained call-reason taxonomy. High-Load Budgeted Categorization of Customer Care Calls describes a deployed encoder–LLM cascade that routes customer-call summaries into fine-grained, long-tailed categories; this paper asks whether the labels a classifier of that kind learns from deserve trust. Disagreement among labelers here collapses onto a few pairs of semantically adjacent categories, which points back at the category definitions. Clusters Are Proposals comes at the taxonomy from the other end: it discovers finer subcategories inside broad call groups, gives each candidate a written definition, and merges two candidates when an LLM confirms they describe the same subcategory.

## What this does not show

The evidence comes from one task and one corpus, so the halo's size on other label spaces is untested; the paper conjectures, without testing, that it may appear in medical coding, moderation taxonomies and LLM juries. Within the panel, two vendor lines contribute sibling models (Anthropic and OpenAI), and the effect is most directly evidenced for the Anthropic pair. Same-family and cross-family sources are not matched on capability, so capability proximity rather than lineage may drive part of the halo; the paper flags this confound and does not resolve it. The 20 exploratory contrasts share the same calls and overlapping models and prompts, so they corroborate the headline rather than replicate it. The effective-labeler count is an approximation, not an exact count. The randomized construction's benefit is directional and measured against gold, with every bootstrap interval including zero at 924 calls. Nor does the paper show that a family-diverse panel yields a more accurate downstream classifier, or settle how much of the general silver–gold divergence is shared bias and how much is models being right where humans were wrong. The call data cannot be released.

## Blog post

[An LLM labeler scores higher when its sibling wrote the answer key](https://telegrapher.ai/blog/same-family-halo.md)

## Cite

```bibtex
@misc{arbuzov2026the,
  title         = {The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver Labeling},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Dave, Karan and Lao, Vincent and Jain, Navita},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/same-family-halo/}
}
```
