telegrapher

Rare call issues can get lost before an LLM ever sees them

Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized TextThe paper: summary, reader and PDF

A big telecom carrier sorts tens of thousands of customer calls a day into broad buckets such as Make a Payment or Activate a Device. Those buckets serve routing and staffing, not the question an operations team asks: of the many reasons people call about payments, which should be fixed first? One “payment issues” cluster can hold autopay failures, surprise charges and app errors, each needing its own fix.

The issues worth catching early tend to be the hardest to spot. An app update that broke payments last week, or eSIM setup breaking on one phone model, might come to thirty calls in three thousand. By the time it is loud enough to notice, customers and agents have been paying for it for weeks.

In the paper we trace the step where LLM discovery pipelines lose issues that small, and change it.

Most samples show a 1% problem fewer than three times

Recent LLM topic discovery shares a sound template we call propose-then-assign: an LLM proposes categories with written definitions, then every document is labeled against them. We keep it.

The weak link is the proposal step: the LLM sees a random sample and cannot name what the sample missed. A random sample of 200 documents misses a 1% subcategory entirely about 13% of the time; in about 68% of samples it turns up fewer than three times, too few for any proposer to see a pattern.

Later stages make it worse. One system prunes candidates generated too rarely; another folds fine distinctions into a few coarse labels; a third drops clusters below a minimum size. Each choice is reasonable in the system that made it. Sampling can keep a small group from the proposer altogether, and the later rules can remove what little gets through.

Ordinary clustering does not rescue it. In Make a Payment every call points roughly the same way in embedding space, with subcategory centroids typically above 0.85 cosine similarity, so a rare subcategory is a slight bump on a crowded surface. One density-clustering pass tuned to the main subcategories folds it into a neighbor.

Ask the clustering to be exhaustive, not right

Take one category. We run UMAP and HDBSCAN on its call embeddings and split any cluster larger than 50 calls again, with no fixed depth, until each piece is at or below that size or stops splitting. This deliberately produces too many pieces. In the paper’s illustration, calls about updating an expired card and calls about replacing one land in separate clusters. Every piece goes to the LLM as a candidate.

That is what we mean by clusters as proposals: a cluster is a candidate to be read and judged, not an answer. Underneath sits an asymmetry. A redundant proposal costs one merge later. A proposal that never forms is lost for good, because nothing downstream goes looking for it.

So the judgment goes into the merge, and raw similarity cannot make it; inside one category, duplicates and mere neighbors look alike. We first subtract the category centroid, removing the direction the whole category shares, and compare what is left: the direction in which each cluster specializes. For each close pair, an LLM reads 20 sampled calls from each cluster and confirms or rejects the merge. The two expired-card clusters from the illustration are joined this way.

The LLM then gives each surviving candidate a name and a one-sentence definition. Each call is matched on its own against those definitions rather than inheriting its cluster, and stays unassigned when none fits.

Ten categories, three methods

We ran this on the ten highest-volume categories of a North American telecom provider’s call-center corpus, each call already reduced upstream to a one-sentence reason for calling. Flat clustering runs once with the same embeddings and tuned settings. LLM-first discovery has a frontier model write definitions straight from document samples, then assigns every call against them.

Start with the cost. Geometry flagged 80 merge candidates among 244 proposals; the LLM confirmed 10, a redundancy rate of 4.1%. It judged the other 70 to describe different subcategories, so by its reading geometry alone would have over-merged. The rate also reflects a setting we chose: the merge prompt asks for clear duplicates and keeps borderline distinctions apart, a dial a deployment can turn the other way.

Method Subcategories Coherence Calls unassigned
Flat clustering 111 0.792 15.3%
LLM-first discovery 203 0.758 0.2%
Clusters as proposals 234 0.835 28.6%

Coherence is the mean pairwise cosine similarity within a subcategory. Against LLM-first it is higher in all ten categories, including the three where LLM-first finds more topics. In Make a Payment, flat clustering returns 7 groups, LLM-first 18, ours 24. In eight of the ten categories, at least one discovered subcategory holds under 2% of the category’s calls.

Because coherence lives in the clustering’s own embedding space, an LLM judge also rated each topic with at least ten calls for specificity and actionability. Against LLM-first the gain is modest, a fifth to a quarter of a point on a five-point scale, yet steady enough to be significant on both measures (p=0.002). Against flat clustering the overall difference is not significant (p=0.16). Where flat clustering finds two to four large groups, actionability rises by over half a point on average; in four other categories, where it already finds many, we score lower. We treat that split as exploratory.

Finer topics, fewer calls covered

Service Not Restored is where the trade-off bites. Flat clustering finds three groups there, the largest over three thousand calls. We find 23 topics, which the judge rates more actionable on average; the share of calls left unassigned climbs from 1% to 37%.

That is the intent: a call that does not clearly match any fine-grained definition is meant to stay unassigned rather than be forced into a poor fit. A team hunting for an issue nobody has named may accept that; one that needs every call accounted for will not.

Spend the LLM on the merge, not the sample

With a coarse taxonomy and small issues to find, the lever is what the LLM gets shown. Over-generating is cheap when duplicates are rare and each costs one checked merge.

High-Load Budgeted Categorization of Customer Care Calls studies a deployed cascade that routes customer-call summaries into fine-grained, long-tailed categories; on escalated calls, where the LLM picks from a shortlist, grounding each candidate in a synthesized definition lifts conditional accuracy. This paper covers the step before, finding finer categories and defining them.

Recall is the open question

All of this comes from one enterprise corpus with no reference taxonomy. We can say the topics are more numerous and, on average, more coherent; we cannot say they are the true subcategories, or that none were missed. The public test, CLINC150 and BANKING77 with their fine intent labels hidden, is named but not yet run.

Coherent does not mean distinct. A pairwise check found substantial overlap among nearest-neighbor topics, so counts of highly rated topics overstate how many distinct issues were found.

Coverage is uneven, and we make no comparison at matched coverage. Coherence also feeds the tuning objective in the reported run, which favors the embedding-based methods. We have not reported the sub-2% count for the baselines.

The judge, Claude Sonnet 4.6, touches both sides: it runs the LLM-first baseline, which could favor that baseline, and it generated the paraphrases our embeddings average in, which could favor us. The Same-Family Halo is why either link matters: agreement among models can reflect shared labeling preferences rather than independent confirmation. And the ratings measure LLM-judged usefulness, not business outcomes.

Recursion has limits too. A group below HDBSCAN’s minimum cluster size, searched from 20 to 200, cannot become its own proposal. Nor does splitting always isolate one problem: a Make a Payment topic of 197 calls still mixes four causes.

What we most want to know is whether recursion recovers the small subcategories a known taxonomy says are there, the ones a random sample might have skipped. The method needs no change to face that test.

Read the paperClusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized TextICLR 2027 submission, September 2026