Over-splitting call categories finds 234 subcategories where flat clustering finds 111
Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text
Navita Jain, Mikhail L Arbuzov, Dmitry Dimov, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Karan Dave, Sisong Bei
ICLR 2027 submission, September 2026
Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not.
What we did and found
Recent LLM pipelines for topic discovery share a sound template, propose-then-assign: an LLM proposes categories with written definitions from a sample of documents, and every document is then labeled against them. The sample is where things go wrong. A subcategory holding 1% of a category is missing from a random 200-document sample about 13% of the time, and in about 68% of samples it shows up fewer than three times; later stages then prune rare candidates, fold fine distinctions into coarse labels, or drop clusters below a minimum size. Clusters as proposals changes where the candidates come from. UMAP and HDBSCAN split each category, and any cluster larger than the threshold of 50 documents is split again, with no depth cap, until each piece fits or splitting stops making progress. Every leaf reaches the LLM as a candidate. To find duplicates, cluster centroids are compared after the category centroid is subtracted (delta-vector similarity); for each flagged pair, an LLM reads sampled calls from both clusters and confirms or rejects the merge. Survivors get a name and a one-sentence definition. Each call is then assigned against those definitions, or left unassigned when none fits.
Across the 10 highest-volume categories of a production telecommunications corpus, the method finds 234 subcategories. Flat clustering, with the same embeddings and tuning, finds 111; an LLM-first baseline finds 203. Mean coherence is higher than both, and higher than LLM-first in all 10 categories. Over-splitting turned out cheap: geometry flagged 80 merge candidates among 244 proposals, and the LLM confirmed 10, a 4.1% redundancy rate. In 8 of the 10 categories at least one discovered subcategory holds less than 2% of its category's calls. An LLM judge rates the topics more specific and more actionable than LLM-first's (p=0.002). Against flat clustering the picture is conditional: actionability rises by over half a point on average where flat clustering finds 2–4 groups, and falls in the four categories where it already finds many. The cost shows up in coverage, with 28.6% of calls unassigned against 15.3% for flat clustering and 0.2% for LLM-first.
Key numbers
| Subcategories discoveredacross 10 call categories, against 111 from flat clustering and 203 from LLM-first discovery | 234 |
| Redundancy rate10 LLM-confirmed merges among 244 proposals; centroid-subtracted geometry had flagged 80 candidates | 4.1% |
| Mean topic coherencemean pairwise cosine similarity within a topic, against 0.792 for flat clustering and 0.758 for LLM-first | 0.835 |
| Mean actionability (LLM judge)1–5 scale, category macro-average, against 3.58 for both flat clustering and LLM-first | 3.78 |
| Calls left unassignedagainst 15.3% for flat clustering and 0.2% for LLM-first | 28.6% |
What this does not show
Everything rests on one enterprise corpus with no reference taxonomy, so the results speak to coherence and granularity, not to recall of true subcategories or semantic completeness; public benchmarks with known fine intents (CLINC150, BANKING77) are named as the test but not run. Coherence and separation are computed in the same embedding space the clustering uses, and in the reported run they also enter the tuning objective, which favors the embedding-based methods over LLM-first. The 4.1% redundancy rate counts the merges the LLM accepted, 10 against 244 proposals, under a prompt set to keep borderline distinctions; a separate pairwise check found substantial overlap among nearest-neighbor topics, and the paper warns that counts of highly rated topics overstate distinct actionable discoveries. Coverage differs sharply, 28.6% unassigned against 0.2% for LLM-first, and no comparison is made at matched coverage. The judge, Claude Sonnet 4.6, is also the LLM-first baseline's model, and its ratings measure LLM-judged usefulness rather than business outcomes. Against flat clustering the judged gain is not significant (p=0.16) and reverses in four categories; coherence falls in at least one of them, Update Account Details. The count of 8 categories with a sub-2% subcategory is not reported for the baselines. A minimum cluster size still bounds how small a proposal can be, and recursion does not always isolate one problem: a Make a Payment topic of 197 calls still mixes several causes.
The authors' abstract
Enterprise call-center analytics often begins with broad topic groups, while the finer distinctions needed for actionable analysis remain hidden. Discovering these distinctions is challenging when interactions within each group are semantically similar, category frequencies are highly uneven, and the number of categories is unknown. The challenge is sharpest when meaningful subcategories are rare: an infrequent but important customer issue may represent only 1% of a broad category’s traffic. Recent LLM pipelines discover categories by proposing candidate definitions from a sample of documents and then assigning every document against them. We argue that these pipelines lose hidden subcategories before an LLM ever sees them: proposals are drawn from random samples, and small candidates are later pruned, absorbed into coarse labels, or removed by minimum-size rules. We introduce clusters as proposals, which changes where proposals come from. Each category is recursively over-fragmented into size-bounded clusters, so every dense region, however small, reaches the LLM as a candidate subcategory with a written definition. Candidates are merged only when an LLM confirms, from their definitions and supporting texts, that they describe the same subcategory, and every document is then labeled against the final definitions. The design rests on an asymmetry: a redundant proposal costs one merge, but a missing proposal cannot be recovered later. On 10 high-volume categories of a production telecommunications call-center corpus, cluster proposals yield 234 subcategories, compared with 111 from flat clustering and 203 from LLM-first discovery, with higher within-subcategory coherence than both baselines in all 10 categories. In 8 of those categories, at least one discovered subcategory holds less than 2% of that category’s calls—the hidden issues the method is designed to surface.
More in Conversation analytics
Cite
@misc{jain2026clusters,
title = {Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text},
author = {Jain, Navita and Arbuzov, Mikhail L and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Dave, Karan and Bei, Sisong},
year = {2026},
note = {ICLR 2027 submission},
url = {https://telegrapher.ai/research/clusters-are-proposals/}
}Builds on
- Pham et al. (2024). TopicGPT: A prompt-based topic modeling framework.
- Wan et al. (2024). TnT-LLM: Text mining at scale with large language models.
- Grootendorst (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure.
- Campello et al. (2013). Density-based clustering based on hierarchical density estimates.
- Tamkin et al. (2024). Clio: Privacy-preserving insights into real-world AI use.