telegrapher

An LLM labeler scores higher when its sibling wrote the answer key

The Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver LabelingThe paper: summary, reader and PDF

A customer calls to ask for a few more days to pay a bill. In a taxonomy with hundreds of fine-grained call reasons, that call could land under a payment or under a payment arrangement or extension, and a label space crowded with near neighbors like these is hard for humans and models alike. At tens of millions of conversations, no annotation team can label every call. The common move is to have a language model write the labels and train a classifier on them; in the pipeline we studied, roughly 76,000 model-labeled interactions do that training.

Those model-written labels are called silver labels, and the usual way to check them is agreement. Ask a second model. If it picks the same category, read the match as confirmation; if not, flag the call as hard. The logic borrows from the old idea that a crowd of voters beats any single voter, and it carries the old premise along: the voters have to be independent. Language models trained on overlapping data toward overlapping objectives are a strange crowd to assume that of.

In the paper we test that premise on a production call-categorization pipeline. It does not hold, and the way it breaks is predictable.

Human labels make a broken ruler

The first instinct is to check silver against human gold. We have a holdout of 924 human-labeled calls, and every labeler we ran agrees more with other models’ labels than with the humans’. That looks like shared model bias. It might be. But a manual review of the disagreements suggests part of the gap is gold error, calls where the human label is plainly wrong and the model’s is right. A model that “beats” gold could share a blind spot with its peers, or simply be right more often than the annotator. Against gold, the two cases look the same.

So we took gold out of the bias question.

Hold the labeler fixed, change who wrote the key

Take one labeler, Haiku. Its predictions on the 924 calls never change. What changes is the answer key we grade them against. First we score Haiku against the labels Sonnet wrote for the same calls; then against Gemini’s, GPT-5.5’s and Grok’s. Haiku’s competence is identical in all four comparisons and no human label enters any of them, so any difference in agreement has to come from who wrote the key.

Sonnet and Haiku are siblings, two models from one vendor line. Gemini, GPT-5.5 and Grok come from three other families. Two leave-out rules keep the comparison honest. A labeler is never graded against its own output, which would score 100% by construction, and when the key is a vote across a panel, the labeler is removed from that vote first.

That within-labeler difference is the gold-free audit. When the sibling’s key wins, the excess agreement is what we call the same-family halo.

A sibling’s key flatters the labeler

Silver answer key Family, relative to Haiku Haiku’s agreement
Sonnet-4.6 same 0.838
Gemini-2.5-Pro cross 0.799
GPT-5.5 cross 0.799
Grok-4.6 cross 0.801

Haiku agrees with its sibling’s labels 3.7 to 3.9 points more than with any of the three cross-family keys, and each gap survives a Bonferroni correction. The three outside keys, from three different vendors, land within two thousandths of each other. The sibling sits apart.

Haiku is the cleanest case, not the whole story. Across 20 exploratory contrasts, crossing four frontier answer keys with five prompts, 17 are significant and 16 of those point the halo’s way, by as much as 5.4 points. The one significant exception is a single Grok prompt under which same-family labelers grade each other more harshly. The 20 share calls and overlapping models, so they corroborate the headline rather than replicate it twenty times.

A second, different statistic agrees. Sonnet and Haiku fail on the same calls more often than Sonnet does with models from other families.

Families decorrelate; prompts set quality

Correlated errors reach beyond siblings, though. When two labelers tend to fail on the same calls, the second adds less than a full vote, and a design-effect formula turns a panel’s average error correlation into an effective number of independent labelers. By that count our full grid of 48 labeler configurations, cross-family pairs included, behaves like fewer than two. Five prompts on Sonnet behave like roughly one.

Which lever buys more independence: a different model family, or a different prompt? Families, narrowly. Every panel of five prompts on one model came out less independent than every panel of one prompt across six families, but the margin is thin, and none of these panels is worth two independent labelers. Prompt variants of one model tend to miss the same calls.

That does not make prompts unimportant. A prompt barely moves a strong model and moves a weak one a great deal: across five prompts, gold accuracy swings by 3.2 points on our strongest model and by 18.0 on the weakest.

Where labelers do disagree, the disagreements are not scattered. They collapse onto a few clusters of semantically adjacent categories (the largest links Customer Prospecting with Inquiries about Services), which suggests that part of the “error” lives in overlapping definitions, or in calls too thin to tell two categories apart, rather than in the models.

What to do with a halo

When you build a panel to vet silver labels, spend its diversity on model families rather than on prompt variants of one model. Do not count agreement between siblings as stronger confirmation than agreement across families. And do not let one fixed source write the whole answer key, because whichever model authors it hands a systematic advantage to its own family, and the system trained downstream inherits the tilt.

For that last point we propose combinatorial silver-label construction. Instead of one model and one prompt labeling the training corpus, each call gets its silver label from a (model, prompt) pair drawn at random from the pool. Over a corpus, no single family or prompt writes enough of the key to be systematically flattered. The prompt is randomized for the quality reason above: a corpus written under one prompt inherits that prompt’s quality profile. The call count does not grow. Each label still costs exactly one inference call, the same as a fixed source and a fraction of what a voting panel pays, although per-token prices differ and the dollar cost follows the pool’s average.

Does it help? Directionally. Every diversified construction we tried reduced how unevenly the families are flattered, compared with a single fixed source, and the lowest single value came from randomizing both the model family and the prompt for each call. What randomization cannot do is remove the errors two siblings make together. It spreads the flattery evenly; the shared mistakes are stubborn, and they survive any vote and any way of building the key.

Before and after the labels

Clusters Are Proposals works upstream, on where finer call categories come from: it finds candidate subcategories inside broad topic groups and gives each a written definition. High-Load Budgeted Categorization works downstream. It describes a deployed cascade in which a fine-tuned encoder answers the calls it can and escalates the rest to an LLM; this paper is about the labels a classifier of that kind learns from.

What we have not shown

Everything here comes from one task and one corpus, so the halo’s size on other label spaces is untested. Inside our panel, only two vendor lines contribute sibling models, Anthropic and OpenAI, and the effect is most directly evidenced for the Anthropic pair.

The harder confound is capability. Our same-family and cross-family sources are not matched on strength, so part of what looks like lineage could be two models of similar ability making similar calls. Separating the two needs strength-matched sources, which we have not run.

The rest are smaller. The randomized construction’s benefit is a suggestion, not a result: at 924 calls every reduction’s confidence interval includes zero, and the spread it reduces is measured against gold. The effective-labeler count is an approximation meant for intuition, not an exact count. Nor have we shown that a family-diverse panel trains a more accurate downstream classifier, or settled, with gold this imperfect, how much of the general gap between model and human labels is bias and how much is competence.

We conjecture, without having tested it, that the same source dependence may turn up wherever LLM consensus certifies labels over a large, fine-grained label space: in medical coding, in content-moderation taxonomies, in LLM juries that score other models. The most direct test is a public consumer-complaint or intent-classification corpus. Because the audit never consults gold, that test needs no new human annotation, just several model families labeling the same items.

Read the paperThe Same-Family Halo: A Gold-Free Audit of Source-Dependent Agreement in LLM Silver LabelingICLR 2027 submission, September 2026