---
type: paper
slug: clusters-are-proposals
title: 'Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized
  Text'
authors:
- Navita Jain
- Mikhail L Arbuzov
- Dmitry Dimov
- Evgeniya Dontsova
- Yaodong Hu
- Vincent Lao
- Karan Dave
- Sisong Bei
date: '2026-09-17'
status: ICLR 2027 submission
line: Conversation analytics
pages: 20
html: https://telegrapher.ai/research/clusters-are-proposals/
pdf: https://telegrapher.ai/papers/clusters-are-proposals/clusters-are-proposals.pdf
reader: https://telegrapher.ai/research/clusters-are-proposals/read/
json: https://telegrapher.ai/api/papers/clusters-are-proposals.json
openreview: https://openreview.net/forum?id=NoOxt4FXen
---

# Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text

## Paper gist

- **Claim:** Over-splitting call categories finds 234 subcategories where flat clustering finds 111
- **TL;DR:** Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not.
- **Method:** On single-sentence, LLM-distilled statements of each caller's primary reason for calling, from the 10 highest-volume categories of a production telecommunications call-center corpus, embedded with GTE-Large, recursive UMAP and HDBSCAN proposals with centroid-subtracted, LLM-confirmed merging and LLM assignment against written definitions are compared with single-pass flat clustering and an LLM-first baseline (Claude Sonnet 4.6) on embedding coherence and separation, unassigned rate, and LLM-judged specificity and actionability.
- **Key result:** Subcategories discovered: 234; Redundancy rate: 4.1%; Mean topic coherence: 0.835
- **Why it matters:** The readers this serves already have a taxonomy.
- **Limits:** Everything rests on one enterprise corpus with no reference taxonomy, so the results speak to coherence and granularity, not to recall of true subcategories or semantic completeness; public benchmarks with known fine intents (CLINC150, BANKING77) are named as the test but not run.
- **Status:** ICLR 2027 submission, September 2026
- **Read:** reader /research/clusters-are-proposals/read/, PDF /papers/clusters-are-proposals/clusters-are-proposals.pdf

## Abstract

Enterprise call-center analytics often begins with broad topic groups, while the finer distinctions needed for actionable analysis remain hidden. Discovering these distinctions is challenging when interactions within each group are semantically similar, category frequencies are highly uneven, and the number of categories is unknown. The challenge is sharpest when meaningful subcategories are rare: an infrequent but important customer issue may represent only 1% of a broad category’s traffic. Recent LLM pipelines discover categories by proposing candidate definitions from a sample of documents and then assigning every document against them. We argue that these pipelines lose hidden subcategories before an LLM ever sees them: proposals are drawn from random samples, and small candidates are later pruned, absorbed into coarse labels, or removed by minimum-size rules. We introduce clusters as proposals, which changes where proposals come from. Each category is recursively over-fragmented into size-bounded clusters, so every dense region, however small, reaches the LLM as a candidate subcategory with a written definition. Candidates are merged only when an LLM confirms, from their definitions and supporting texts, that they describe the same subcategory, and every document is then labeled against the final definitions. The design rests on an asymmetry: a redundant proposal costs one merge, but a missing proposal cannot be recovered later. On 10 high-volume categories of a production telecommunications call-center corpus, cluster proposals yield 234 subcategories, compared with 111 from flat clustering and 203 from LLM-first discovery, with higher within-subcategory coherence than both baselines in all 10 categories. In 8 of those categories, at least one discovered subcategory holds less than 2% of that category’s calls—the hidden issues the method is designed to surface.

## What we did and found

Recent LLM pipelines for topic discovery share a sound template, propose-then-assign: an LLM proposes categories with written definitions from a sample of documents, and every document is then labeled against them. The sample is where things go wrong. A subcategory holding 1% of a category is missing from a random 200-document sample about 13% of the time, and in about 68% of samples it shows up fewer than three times; later stages then prune rare candidates, fold fine distinctions into coarse labels, or drop clusters below a minimum size. Clusters as proposals changes where the candidates come from. UMAP and HDBSCAN split each category, and any cluster larger than the threshold of 50 documents is split again, with no depth cap, until each piece fits or splitting stops making progress. Every leaf reaches the LLM as a candidate. To find duplicates, cluster centroids are compared after the category centroid is subtracted (delta-vector similarity); for each flagged pair, an LLM reads sampled calls from both clusters and confirms or rejects the merge. Survivors get a name and a one-sentence definition. Each call is then assigned against those definitions, or left unassigned when none fits.

Across the 10 highest-volume categories of a production telecommunications corpus, the method finds 234 subcategories. Flat clustering, with the same embeddings and tuning, finds 111; an LLM-first baseline finds 203. Mean coherence is higher than both, and higher than LLM-first in all 10 categories. Over-splitting turned out cheap: geometry flagged 80 merge candidates among 244 proposals, and the LLM confirmed 10, a 4.1% redundancy rate. In 8 of the 10 categories at least one discovered subcategory holds less than 2% of its category's calls. An LLM judge rates the topics more specific and more actionable than LLM-first's (p=0.002). Against flat clustering the picture is conditional: actionability rises by over half a point on average where flat clustering finds 2–4 groups, and falls in the four categories where it already finds many. The cost shows up in coverage, with 28.6% of calls unassigned against 15.3% for flat clustering and 0.2% for LLM-first.

## Key numbers

| Measure | Value |
|---|---|
| Subcategories discovered | 234 |
| Redundancy rate | 4.1% |
| Mean topic coherence | 0.835 |
| Mean actionability (LLM judge) | 3.78 |
| Calls left unassigned | 28.6% |

## Why it matters

The readers this serves already have a taxonomy. Their categories route calls and set staffing well enough; the trouble sits inside them. A payment-app regression shipped last week, or eSIM provisioning failing on one device model, may come to thirty calls out of three thousand, and by the time it is frequent enough to notice it has been costing customers and agents for weeks. The paper puts the loss before any LLM reads a call. Each pruning rule in earlier pipelines protects something reasonable, whether label quality, cost or privacy; stacked together, they can remove a small group before the proposer ever sees it. Nothing about the LLM changes here. What changes is what it gets shown.

The paper argues, without testing it, that the design applies wherever documents arrive pre-sorted into categories and finer ones are needed. The rule underneath is simple: if a later stage can delete a redundant candidate but nothing downstream can recover a missing one, generate too many and spend judgment on the merge. Geometry cannot make that judgment alone. Inside a narrow category, raw cosine similarity between cluster centroids stays above 0.85, so duplicates and mere neighbors look alike; subtracting the category centroid leaves the direction in which each cluster specializes, and that is what gets compared. Even so, of 80 flagged candidates the LLM judged 70 to describe different subcategories.

The work sits in the group's conversation-analytics line. High-Load Budgeted Categorization of Customer Care Calls routes customer-call summaries into fine-grained, long-tailed categories and finds that, on the escalated calls the LLM sees with a shortlist of labels, grounding each candidate label in a synthesized definition lifts conditional accuracy; this paper comes at the same problem from the other end, discovering finer subcategories inside broad ones and writing a definition for each. The Same-Family Halo adds a caution that applies directly. Agreement among models can reflect shared labeling preferences rather than independent confirmation, and the judge in this evaluation is the same model that runs the LLM-first baseline.

## What this does not show

Everything rests on one enterprise corpus with no reference taxonomy, so the results speak to coherence and granularity, not to recall of true subcategories or semantic completeness; public benchmarks with known fine intents (CLINC150, BANKING77) are named as the test but not run. Coherence and separation are computed in the same embedding space the clustering uses, and in the reported run they also enter the tuning objective, which favors the embedding-based methods over LLM-first. The 4.1% redundancy rate counts the merges the LLM accepted, 10 against 244 proposals, under a prompt set to keep borderline distinctions; a separate pairwise check found substantial overlap among nearest-neighbor topics, and the paper warns that counts of highly rated topics overstate distinct actionable discoveries. Coverage differs sharply, 28.6% unassigned against 0.2% for LLM-first, and no comparison is made at matched coverage. The judge, Claude Sonnet 4.6, is also the LLM-first baseline's model, and its ratings measure LLM-judged usefulness rather than business outcomes. Against flat clustering the judged gain is not significant (p=0.16) and reverses in four categories; coherence falls in at least one of them, Update Account Details. The count of 8 categories with a sub-2% subcategory is not reported for the baselines. A minimum cluster size still bounds how small a proposal can be, and recursion does not always isolate one problem: a Make a Payment topic of 197 calls still mixes several causes.

## Blog post

[Rare call issues can get lost before an LLM ever sees them](https://telegrapher.ai/blog/clusters-are-proposals.md)

## Cite

```bibtex
@misc{jain2026clusters,
  title         = {Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text},
  author        = {Jain, Navita and Arbuzov, Mikhail L and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Dave, Karan and Bei, Sisong},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/clusters-are-proposals/}
}
```
