Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text
Back to the paper page. ICLR 2027 submission, September 2026.
All 20 pages are shown below.
Text of page 1
Under review as a conference paper at ICLR 2027 C LUSTERS A RE P ROPOSALS : D ISCOVERING H IDDEN S UBCATEGORIES IN P RE -C ATEGORIZED T EXT Anonymous authors Paper under double-blind review A BSTRACT Enterprise call-center analytics often begins with broad topic groups, while the finer distinctions needed for actionable analysis remain hidden. Discovering these distinctions is challenging when interactions within each group are semantically similar, category frequencies are highly uneven, and the number of categories is unknown. The challenge is sharpest when meaningful subcategories are rare: an infrequent but important customer issue may represent only 1% of a broad category’s traffic. Recent LLM pipelines discover categories by proposing candidate definitions from a sample of documents and then assigning every document against them. We argue that these pipelines lose hidden subcategories before an LLM ever sees them: proposals are drawn from random samples, and small candidates are later pruned, absorbed into coarse labels, or removed by minimum-size rules. We introduce clusters as proposals, which changes where proposals come from. Each category is recursively over-fragmented into size-bounded clusters, so every dense region, however small, reaches the LLM as a candidate subcategory with a written definition. Candidates are merged only when an LLM confirms, from their definitions and supporting texts, that they describe the same subcategory, and every document is then labeled against the final definitions. The design rests on an asymmetry: a redundant proposal costs one merge, but a missing proposal cannot be recovered later. On 10 high-volume categories of a production telecommunications call-center corpus, cluster proposals yield 234 subcategories, compared with 111 from flat clustering and 203 from LLM-first discovery, with higher within-subcategory coherence than both baselines in all 10 categories. In 8 of those categories, at least one discovered subcategory holds less than 2% of that category’s calls—the hidden issues the method is designed to surface. 1 I NTRODUCTION A large telecommunications provider routes tens of thousands of customer calls a day into broad categories such as Make a Payment or Activate a Device. The categories work well for routing and staffing, but are too broad to guide operational decisions—there are twenty different reasons customers call about payments; which ones should we act on first? A single broad cluster labeled “payment issues” may combine calls about autopay failures, unexpected charges, and payment-app errors—distinct problems that require different responses. Manual review cannot answer it at scale, because human reviewers evaluate far less than 1% of calls (Stepanov et al., 2015). The answer that matters most is also often the most hidden. A payment-app regression introduced last week, or eSIM provisioning failing on one device model, may account for thirty calls out of three thousand. By the time such an issue is frequent enough to be obvious, it has been costing customers and agents for weeks. Large language models have made this kind of discovery practical, and recent work has converged on a sound template that we call propose-then-assign (Pham et al., 2024; Wan et al., 2024; Wang et al., 2023; Lam et al., 2024). An LLM reads documents and proposes candidate categories with natural-language definitions; the candidates are refined; and every document is then labeled against the final definitions rather than by where it happens to fall in an embedding space. Separating 1 Reviewers: please read the Reviewer Guidelines (iclr.cc/Conferences/2027/ReviewerGuidelines) and the AI Policy for Reviewers (iclr.cc/Conferences/2027/AIPolicyForReviewers). If you used AI to expand, edit, or polish your review, please provide the input text to the LLM. Better still, consider skipping the LLM and submitting your original text: we, and the authors, are much more interested in your unedited thoughts than in what an LLM has to say. AI-assisted or not, you are putting your name and reputation behind your review: LLM-generated falsehoods, hallucinations or misrepresentations are subject to disciplinary action, which may include desk-rejecting all papers you have authored.
Text of page 2
Under review as a conference paper at ICLR 2027 Figure 1: Our approach on a single call category. Flat clustering (left) returns a few broad groups. Our approach (right): Recursive over-fragmentation turns every dense region into a proposal → category-centered geometric candidates and LLM merges redundant proposals (2 pairs identified in the over-fragmented clusters) → LLM defines each proposal and documents are labeled against the final definitions. discovery from assignment makes the output interpretable and every label auditable. We adopt this template. Its weak point is the first step. An LLM can only propose what it is shown, and in current pipelines it is shown a random sample of documents. The arithmetic is unforgiving. A subcategory that holds 1% of a category has a 13% chance of being entirely absent from a random sample of 200 documents, and a 68% chance of appearing fewer than three times—too few for any proposer to recognize a pattern. Later stages then compound the loss. Candidates generated too rarely are pruned as noise (Pham et al., 2024), fine distinctions are folded into a small set of coarse labels (Wan et al., 2024), and clusters below a minimum size are dropped outright (Tamkin et al., 2024). Each of these choices is reasonable where it was made: it protects label quality, cost, or privacy. Together, they mean that hidden subcategories are removed before an LLM ever sees them. The problem is sharpest in exactly our setting, where discovery runs inside a category that is already semantically narrow. Every document in Make a Payment shares a strong common direction in embedding space, and the cosine similarity between any two subcategory centroids is typically above 0.85. A hidden subcategory is therefore not a separate island but a small bump on a dense surface. A random sample rarely lands on it, and a single pass of density clustering tuned for the main subcategories absorbs it into a larger neighbor. We propose to change where proposals come from. Our central idea is to treat clusters as proposals rather than answers. Instead of asking a clustering to be correct, we ask it to be exhaustive. Each category is recursively split without a fixed depth cap until each branch meets the size threshold or further splitting makes no progress, and every resulting cluster is shown to the LLM, which writes a definition for it. This deliberately produces too many candidates, and that is the point, because the two possible errors are not symmetric: a redundant proposal costs one merge, while a missing proposal cannot be recovered by any later stage. Consolidation removes the redundancy. Candidates are compared after subtracting the category centroid, which removes the shared direction that makes every subcategory look alike; an LLM confirms each merge from the candidates’ definitions and supporting texts; and every document is labeled against the surviving definitions. Figure 1 shows the contrast on a single category: in Make a Payment, flat clustering returns 7 broad groups, while our approach identifies 24 distinct subcategories; even the LLM-first baseline, which has access to the same frontier model, recovers only 18. We study the approach on ten high-volume categories of a production call-center corpus without a reference taxonomy. Our contributions: 1. Diagnosis. Standard clustering metrics are structurally blind to rare subcategories: a latent topic can be absorbed into a larger neighbor without moving any clustering score, and sampling-based proposers miss small groups before an LLM ever sees them (Appendix C.1). 2. Method. Clusters as proposals: every cluster above a size threshold is recursively split regardless of its coherence, so every dense region—however small—reaches the LLM as a candidate with a written definition. Consolidation is governed by geometric redundancy 2
Text of page 3
Under review as a conference paper at ICLR 2027 Table 1: How existing discovery methods generate candidates and what happens to a subcategory absorbed by a broader group. Coarse: whether the method runs inside existing coarse labels without needing labeled examples of the fine classes. Method TopicGPT TnT-LLM GoalEx LLooM Ours assigned Fate of a small subcategory / One density clustering of the Cluster membership corpus Clustering with LLM-refined Cluster membership embeddings Viswanathan et k-means, k given Membership, LLMal. corrected Clio k-means over conversation sum- Cluster membership maries DeepAligned k-means; K estimated Cluster membership Documents by BERTopic Top2Vec ClusterLLM Candidates come from Coarse Absorbed by a neighbor or marked noise – No specific mechanism – No specific mechanism – Removed below a minimum size (privacy) – Estimator drops clusters smaller than – N/K ′ Random sample, one document LLM against definitions Pruned if generated too rarely partial at a time Random minibatches of sum- LLM labels, distilled Folded into a small label set – maries classifier Random subsets; K fixed LLM check per docu- Found only if sampled or left uncovered – ment and description One HDBSCAN pass over a LLM scoring against Found only if sampled – sample criteria Recursive, size-triggered clus- LLM against definitions Proposed whenever it forms a dense region ✓ ters of all docs in a category of ≥ m docs (delta-vector projection), not by cluster size, so rare topics are never pruned for being small (§4). 3. Evidence from deployment. On our production corpus, cluster proposals recover 234 subcategories—against 111 from flat clustering and 203 from LLM-first discovery—with higher within-subcategory coherence than both baselines in all ten categories. 2 R ELATED W ORK For hidden subcategories, two questions decide the outcome of any discovery method: where do candidate categories come from, and what happens to the small ones? We organize prior work around these two questions. Table 1 summarizes the answers; per-method details follow. Clusters as the answer. Classical topic models treat the partition as the result. LDA (Blei et al., 2003b) and its hierarchical extension (Blei et al., 2003a) fix topic structure in advance; BERTopic (Grootendorst, 2022) and Top2Vec (Angelov, 2020) cluster embeddings via UMAP (McInnes et al., 2018) and HDBSCAN (Campello et al., 2013), so granularity follows from a single density threshold. LLM-guided variants refine the geometry but keep the partition as the answer: ClusterLLM (Zhang et al., 2023) tunes embeddings on LLM triplet judgments, Viswanathan et al. (2024) correct k-means with pairwise constraints, and Clio (Tamkin et al., 2024) removes groups below a privacy threshold. In every case a hidden subcategory survives only if the partition isolates it. We use the same machinery but give the partition a different job: generating candidates, not answers, tuned to over-fragment rather than to be right. Propose-then-assign with LLMs. The closest line of work separates discovery from assignment. TopicGPT (Pham et al., 2024) shows sampled documents to an LLM one at a time, merging nearduplicates and pruning rare topics. TnT-LLM (Wan et al., 2024) revises a taxonomy over random minibatches of summaries and distills LLM labels into classifiers. GoalEx (Wang et al., 2023) proposes descriptions from random subsets and re-proposes from uncovered documents. LLooM (Lam et al., 2024) clusters LLM-distilled summaries, proposes concepts per cluster, and loops over outliers. We share this separation and merge candidates with LLM confirmation as in TopicGPT, but differ at the proposal stage: all four systems draw proposals from random samples, so a hidden subcategory may never reach the proposer. Our proposals come from an exhaustive partition, so every dense region is proposed by construction rather than by chance. 3
Text of page 4
Under review as a conference paper at ICLR 2027 Clusters as proposals: discover, consolidate, then assign A Broad category Illustrative examples and counts B Recursive proposals C Define and consolidate D Final assignment Payment issues P1: Update expired card a Expired card b Replace expired card a P1 + P2: Merge e Update an expired payment card. P2: Replace expired card a, b Verification unavailable b P3: Retain c Verification unavailable c, e Payment verification is unavailable. P3: Verification unavailable d Paid; service suspended Update expired card Paid; service suspended c d e Cannot verify payment f Ambiguous request Different issues share one coarse label. P4: Retain P4: Paid; service suspended Service remains suspended after payment is recorded. d f Unassigned e changes groups based on its definition. e is initially grouped with P1. LLM confirms equivalence before merging. f matches no definition: abstain. Figure 2: Overview of clusters as proposals. Recursive fragmentation generates candidate subcategories within a coarse category, delta-vector merging proposes candidates. An LLM confirms merges between equivalent proposals and defines each candidate. Documents are then assigned against the consolidated definitions, with abstention when none applies. The illustration shows redundant fragments and distinct issues retained after consolidation; examples and counts are schematic. Category discovery with an unknown number of classes. New intent discovery (Zhang et al., 2021) and generalized category discovery (Vaze et al., 2022) find novel classes in unlabeled data, estimating the class count when unknown. They differ from our setting in two ways: they require labeled examples of known fine-grained classes to define granularity (we have only coarse categories), and their treatment of small classes works against hidden subcategories—DeepAligned removes clusters smaller than N/K ′ by construction. Recent variants handle class imbalance (Zhang et al., 2024; Bai et al., 2023) or exploit coarse-to-fine taxonomies (He et al., 2025), but all still require labeled fine classes. Over-segment, then merge. Deliberately over-segmenting and then consolidating is established in vision: superpixel methods produce small regions that later stages group into objects (Achanta et al., 2012), and deep clustering trains an auxiliary head with excess clusters, discarded at test time (Ji et al., 2019). We bring this principle to LLM-based category discovery, where over-segmentation ensures hidden subcategories reach the proposer and consolidation reasons over natural-language definitions. 3 P ROBLEM F ORMULATION We study fine-grained topic discovery within existing coarse categories. Given documents grouped by a coarse label, the task is to discover subcategories without labeled fine-grained examples or a specified number of topics. The output is a set of subcategory names and definitions, together with document assignments; documents that match no definition may remain unassigned. We call a subcategory hidden relative to a discovery method when that method absorbs a meaningful distinction into a broader topic without representing it separately. Hidden subcategories may be common or rare. The objective is to expose useful distinctions while limiting redundant topics and preserving document coverage. 4 M ETHOD 4.1 S TAGE 1: P ROPOSAL G ENERATION VIA R ECURSIVE O VER -F RAGMENTATION Within each coarse category, we construct an intentionally over-fragmented taxonomy by applying UMAP followed by HDBSCAN and recursively splitting any cluster exceeding a size threshold (τ split ). The threshold is a recursion trigger, not a guaranteed final size bound: with as low as 20, HDBSCAN will nearly always find density sub-structure in a cluster of several hundred points, producing further splits. This is by design: it ensures that hidden subcategories—those a single clustering pass would absorb into larger groups—reach the LLM as distinct candidates. Each result- 4
Text of page 5
Under review as a conference paper at ICLR 2027 ing leaf is a proposal—a candidate sub-topic, not a final assignment. There is no fixed depth cap; each branch stops when its cluster meets the size threshold or further splitting makes no progress. Documents that do not fall into any cluster are considered noise and are kept outside the proposal leaves. Appendix A.1 gives the recursive procedure. 4.2 S TAGE 2: D ELTA -V ECTOR C ONSOLIDATION AND LLM D EFINITION Over-fragmentation produces redundant proposals—fragments of the same subcategory split across recursion levels. To identify candidates for merging, we center proposal centroids µ i , µ j on the coarse-category mean µ c and compute sim δ (i, j) = (µ i − µ c ) ⊤ (µ j − µ c ) . ∥µ i − µ c ∥ ∥µ j − µ c ∥ (1) Subtracting the category centroid removes the shared direction that makes every subcategory look alike; what remains is the direction in which each proposal specializes. Pairs whose centered similarity exceeds a threshold τ δ are flagged as geometric merge candidates, but similarity alone does not establish semantic equivalence. The LLM then processes each candidate pair, reading 20 sampled documents from each cluster, and confirms or rejects the merge. After each pass, we recompute centroids and candidate similarities, stopping when no further merges are accepted or the iteration limit is reached. For each surviving proposal, the LLM generates a short name and a one-sentence definition from up to 20 representative documents selected by proximity to its centroid. Appendix A.2 provides the complete protocol. 4.3 S TAGE 3: A SSIGN D OCUMENTS Each document is independently assigned by the LLM against the consolidated definitions, rather than inheriting its proposal-cluster membership. The LLM selects the best-matching definition or leaves the document unassigned when none applies. Figure 2 illustrates both reassignment and abstention. Implementation. We use similarity-weighted paraphrase-augmented GTE-Large embeddings (1024 dimensions), UMAP with five output dimensions and cosine distance, and a split threshold of 50 with no fixed recursion-depth cap. A shared UMAP/HDBSCAN configuration is selected per category using 50 Optuna TPE trials (Akiba et al., 2019). The reported run scores complete proposal trees for leaf count, noise, balance, coherence, and separation. Appendices A and D provide embedding, scoring, and search details. 5 E XPERIMENTS 5.1 D ATASET AND S ETUP We evaluate on a large-scale customer service corpus from a North American telecommunications provider: 40,000+ daily interactions across 140+ intent categories (“tertiary categories”), with 2,000–4,000 interactions per category. Detailed results are reported on the 10 highest-volume categories, which account for approximately 90% of total interaction volume. The framework is applied identically to all categories; the 10-category subset is chosen for evaluation because all three methods (Flat, LLM-first, and Proposed) were run on these categories, enabling controlled comparison. Core triggers. Raw interactions are full agent–customer conversation transcripts, often thousands of tokens long and dominated by greetings, holds, and troubleshooting back-and-forth. An upstream LLM distills each transcript into a single-sentence core trigger—the customer’s primary reason for calling, stated in their own words (e.g., “My bill is higher than expected,” “I already made a payment but my service is still suspended”). This distillation is critical for two reasons. First, it strips conversational noise so that embedding similarity reflects topical meaning rather than callflow structure. Second, it produces short, semantically dense inputs where standard bag-of-words topic models struggle but embedding-based methods excel. The upstream extraction is a separate system; this paper operates entirely on the distilled core triggers. 5
Text of page 6
Under review as a conference paper at ICLR 2027 Embedding model. GTE-Large (Li et al., 2023) (1024-dimensional, L2-normalized) served via a managed endpoint. Each core trigger is embedded independently. Full experiment-grid and caching details are provided in Appendix B. 5.2 We evaluate cluster quality along four axes: geometric quality of individual topics, geometric distinctness between topics, structural properties of the taxonomy, and operational usefulness judged by an independent LLM. • Topic Coherence measures how tightly a topic’s documents cluster in embedding space: the mean pairwise cosine similarity among all documents assigned to the same topic, averaged over topics. Values range from 0 to 1; higher means the topic groups semantically similar documents. A topic that mixes unrelated issues (e.g., autopay failures and billing disputes in one cluster) scores low. • Topic Separation measures how distinct topics are from one another: 1− mean cosine similarity between all pairs of topic centroids within the same category, so higher values indicate more distinguishable topics. Computed within each category; the overall score reported in all tables is the unweighted mean across categories, applied identically to every method (Flat, LLM-first, and Proposed). A method that produces many near-duplicate topics scores low on separation even if each individual topic is coherent. • Structural metrics characterize the shape of the discovered taxonomy: the number of topics (leaf count), the fraction of documents left unassigned (noise rate), and how evenly documents distribute across topics (balance, measured as the ratio of observed entropy to maximum entropy). These are not quality scores in themselves but diagnostic indicators— a method that assigns every document to one large cluster has perfect noise but trivial leaf count. E VALUATION M ETRICS • LLM-as-Judge evaluates whether discovered topics are operationally useful, not just geometrically clean. For each topic with ≥10 documents, an independent judge (Claude Sonnet 4.6) rates specificity (1–5: is the topic narrow enough to assign to one team?) and actionability (1–5: could an analyst take a concrete action based on this topic?) from a sample of 10 representative documents (Zheng et al., 2023). Each topic is evaluated twice; scores are averaged within categories and reported as unweighted category-macro-averages. Full protocol in Appendix G. Coherence and separation are computed on the same embeddings used for clustering, so they favor methods that align well with the embedding geometry. The LLM-as-judge evaluation is independent of the embedding space and serves as a complementary signal. 5.3 B ASELINES We compare against two baselines that represent the dominant approaches in the literature: 1. Flat clustering: Flat UMAP + HDBSCAN per category with no recursion—the standard BERTopic-style approach (Grootendorst, 2022). This baseline uses the same embeddings and the same Optuna-tuned UMAP/HDBSCAN configuration as the proposed method, but skips recursive splitting: whatever clusters the first pass produces are the final topics. Differences in results therefore isolate the effect of recursive over-fragmentation and consolidation. 2. LLM-first discovery: A frontier LLM (Claude Sonnet 4.6) generates subcategory definitions directly from document samples, then every document is assigned against that taxonomy—the approach of TnT-LLM (Wan et al., 2024) and TopicGPT (Pham et al., 2024). This baseline tests whether a strong LLM can discover subcategories from text alone, without the geometric proposals that our method provides. Appendix F reports ablations of the embedding strategy, merging criterion, and scoring function. 6
Text of page 7
Under review as a conference paper at ICLR 2027 6 R ESULTS 6.1 F LAT C LUSTERING H IDES S UBCATEGORIES Table 2: Three-way comparison across 10 categories. Flat = single-pass UMAP + HDBSCAN; LLM-first = frontier LLM generates subcategory definitions, then assigns documents; Proposed = clusters-as-proposals after consolidation. Method Topics Coherence Separation Noise % Flat LLM-first Proposed (ours) 111 203 234 0.792 0.758 0.835 0.106 0.098 0.116 15.3% 0.2% 28.6% Flat clustering averages 11 topics per category. These groups are internally similar but absorb finer distinctions: a single cluster of 600 payment calls may mix autopay failures, unexpected charges, and app errors. The proposed pipeline discovers 2.1× more topics per category with higher coherence (+5.4%) and separation (+9.4%), at the cost of higher noise (28.6% vs. 15.3%). The elevated noise is by design: documents that do not clearly match any fine-grained proposal are left unassigned rather than forced into an ill-fitting cluster. 6.2 C LUSTER P ROPOSALS VS . LLM-F IRST D ISCOVERY Coherence improves over the LLM-first baseline in all 10 categories (Appendix Table 8). In 3 of those categories the LLM-first baseline finds more topics (e.g., Update Payment Methods: 41 vs. 29), but in each case with lower coherence. Noise, however, rises substantially (28.6% vs. 0.2% for LLM-first)—a coverage–precision tradeoff where documents that do not clearly match any finegrained definition are left unassigned rather than forced into a poor match. 6.3 C ONSOLIDATION I S A D IAL , N OT A F IXED S TEP Over-fragmentation deliberately produces more proposals than the final taxonomy needs. The question is how aggressively to consolidate them. Because the merge decision is made by an LLM reading definitions and sampled documents, the aggressiveness is controlled at the prompt level: a stringent prompt merges only when two proposals are near-identical; a lenient prompt merges whenever they overlap substantially. The same geometric candidates reach the LLM either way—what changes is how readily it confirms. We set the dial to its lenient end: the LLM is prompted to confirm a merge only when both proposals clearly describe the same subcategory, preserving borderline distinctions rather than collapsing them. Table 3 summarizes the result. Table 3: Consolidation statistics on the best 10-category run under lenient merging. Statistic Value Pre-consolidation proposals Merge candidates identified (geometric) Merges accepted (LLM-validated) Post-consolidation topics Redundancy rate 244 80 10 234 4.1% Delta-vector similarity flags 80 of 244 proposals as geometric merge candidates (sim δ ≥ τ δ ). Even under lenient prompting, the LLM confirms only 10—a 4.1% redundancy rate, spanning 4 of 10 categories. The remaining 70 candidates share similar deviation directions but, when the LLM reads their definitions and supporting documents, describe genuinely different subcategories. Geometry alone would over-merge; the LLM gate prevents it. The low redundancy under lenient consolidation means that most over-fragmented proposals are already distinct—the pipeline pays very little for being exhaustive. A more stringent prompt would 7
Text of page 8
Under review as a conference paper at ICLR 2027 reduce topic count further at the cost of collapsing borderline distinctions; we leave that tradeoff to the deployment context. Raw cosine similarity between centroids within a category stays above 0.85 and cannot separate redundant proposals from distinct ones; delta-vector projection strips the shared category direction, making redundancy visible to the LLM in the first place. 6.4 LLM- AS -J UDGE E VALUATION Coherence and separation measure geometric quality but not whether topics are operationally useful. We use an LLM-based actionability protocol (Zheng et al., 2023) with Claude Sonnet 4.6 as the judge. Protocol. For each topic with ≥10 documents, Claude Sonnet 4.6 rates specificity and actionability on 1–5 scales (two independent runs per topic, averaged). Category-macro-averaged scores are the primary statistic; significance is assessed with two-sided Wilcoxon signed-rank tests paired by category (n=10). Full protocol details, eligibility criteria, and call counts are in Appendix G. Results. Table 4 summarizes the evaluation. The proposed method receives higher specificity and actionability than the LLM-first baseline in all ten categories—modest gains on a five-point scale, but consistent enough to reach significance on both measures (two-sided Wilcoxon, p=0.002; Cohen’s d ≈ 0.3). Against flat clustering, the direction is the same but the signal is weaker (p=0.16, actionability), reflecting both the small number of category pairs (n=10) and the fact that gains concentrate in the subset of categories where flat clustering produces very few topics. These ratings measure LLM-judged usefulness, not demonstrated business outcomes. Table 4: LLM-as-judge actionability evaluation. Method Topics Specificity Actionability % Act≥4 Flat LLM-first Proposed (ours) 111 160 / 203 234 3.26 (3.39) 3.19 (3.25) 3.43 (3.46) 3.58 (3.69) 3.58 (3.63) 3.78 (3.80) 48.6 42.5 54.3 Consistent improvement over LLM-first. On both measures, the proposed method exceeds the LLM-first baseline in every category—by roughly a quarter-point on each scale. The 43 LLM-first topics with fewer than 10 documents are not scored; conclusions apply to the 160 eligible topics. Against flat clustering, benefits are conditional. The proposed method improves actionability over flat clustering in six of ten categories (Appendix Table 9). The gains concentrate where flat clustering produces very few, large groups: in the four categories where the baseline finds only 2–4 topics, actionability rises by over half a point on average. Refinement is most valuable precisely where coarse clustering is most likely to absorb distinct subcategories. In the four categories where flat clustering already finds many well-separated groups, the proposed approach scores lower. This supports a category-dependent granularity tradeoff rather than universal superiority. We treat the subgroup analysis as exploratory; it does not establish that recursion caused the gains or identify which specific baseline groups contained the recovered subcategories. Size–actionability relationship. Smaller proposed topics receive higher actionability ratings (Spearman ρ=−0.215, p=0.001); this correlation is weaker for flat clustering and absent for LLMfirst (Appendix G.1). Actionable topic counts. 54.3% of proposed topics are rated actionable (≥ 4), compared to 48.6% for flat and 42.5% for LLM-first; however, pairwise redundancy among highly rated topics means these counts overstate distinct actionable discoveries (Appendix G.2). Failure example. Recursive fragmentation does not always isolate a single problem. In Make A Payment, proposed topic 32 (197 documents) still combines outdated cards, system errors, enrollment logic, and scheduling restrictions—the judge rates it 3/3 (specificity/actionability) and notes it 8
Text of page 9
Under review as a conference paper at ICLR 2027 “would require separate investigations.” This pattern recurs in inherently multi-cause categories (Multi-Symptom, Charges & Fees), where the flat baseline’s broader groups sometimes receive higher actionability scores because the judge views them as coherent enough to route, even though they conceal finer distinctions. 7 D ISCUSSION AND L IMITATIONS Recursive proposals are most useful when existing categories contain broad groups that obscure finer distinctions. Actionability gains concentrate in categories where flat clustering produces few topics; benefits are less consistent where the baseline already finds finer groups. The tradeoff is coverage: the proposed pipeline leaves roughly twice as many documents unassigned as flat clustering (Table 2), so higher coherence does not come free. These comparisons do not establish quality at matched coverage. Detailed granularity, coverage, and size diagnostics are in Appendix C.4. Our evidence comes from a single enterprise corpus with no reference taxonomy, so the results establish geometric quality and granularity rather than recall against ground truth or semantic completeness. While we validate on customer service data, the clusters-as-proposals design applies anywhere documents are pre-sorted into categories and finer sub-topics are needed: scientific papers within venues, support tickets within product areas, legal documents within case types. We name the public test explicitly: CLINC150 (Larson et al., 2019) with its ten released domains given as coarse categories (and 150 fine intents hidden), and the long-tailed CLINC150 and BANKING77 (Casanueva et al., 2020) splits of Zhang et al. (2024), with fine labels hidden. These benchmarks have known ground-truth intents at multiple granularities, making them a direct test of whether recursive proposals recover subcategories that other methods conceal. Re-running on them requires no change to the method. Broader impact: The system is deployed for customer experience analysis, where discovered topics inform product improvements. Privacy is handled upstream: all text is PII-masked before any pipeline component processes it. 8 C ONCLUSION Coarse categories conceal finer distinctions, and current discovery methods—whether clusteringbased or sampling-based—lose hidden subcategories before they can be named: sampling misses small groups, and frequency- or size-based pruning removes them. The loss is asymmetric: a redundant proposal costs one merge, but a missing proposal cannot be recovered by any later stage. We introduced clusters as proposals: each category is recursively over-fragmented so that every dense region reaches the LLM as a candidate, and consolidation—centered on the category’s own geometry and confirmed by the LLM—removes redundancy before documents are labeled against the surviving definitions. On ten categories of a production call-center corpus, recursive proposals yield roughly twice the subcategories of flat clustering and a fifth more than LLM-first discovery, with higher measured coherence than both. Exhaustiveness is cheap: only 4.1% of proposals turn out to be redundant. In 8 of 10 categories, at least one discovered subcategory holds less than 2% of that category’s volume— the hidden issues the method is designed to surface. An independent LLM judge rates the proposed topics more specific and more actionable than the LLM-first baseline in every category (p=0.002, two-sided Wilcoxon). Against flat clustering, the gains concentrate where the baseline produces only 2–4 coarse topics—precisely where hidden subcategories are most likely—and vanish where the baseline already finds finer groups. The method requires no fine-tuning, no labeled examples of fine classes, and no advance specification of the number of subcategories. These results do not establish semantic completeness, topic uniqueness, or realized business impact; comparisons remain subject to differences in document coverage and topic eligibility. AI-U SE S TATEMENT AI assistance was used to revise the paper’s narrative, review methodological choices and interpretations, check consistency among reported quantities, and propose further analyses. Language 9
Text of page 10
Under review as a conference paper at ICLR 2027 models were also used within the experimental pipeline to construct representations and supervision and to make prompted predictions. The authors are responsible for verifying the final text, results, citations, and disclosures against the underlying research records. E THICS S TATEMENT The study analyzes enterprise customer-service calls, and the evaluated label concerns suppression of further product offers. Errors can change which customers continue to receive offers. Aggregate agreement alone does not establish that two systems affect the same individuals, so operational use requires calibration and review at the intended decision threshold. Access to the underlying transcripts is restricted by the deployment’s data agreement. Brand canonicalization and speaker attribution are representation operations; they should not be interpreted as guarantees of anonymization. The example is an attributed agent–customer exchange. Its statements record what the speakers said, including the agent’s offer and the customer’s plans, without independently verifying those claims. R EPRODUCIBILITY S TATEMENT All experiments use fixed random seeds at every pipeline stage. The method is described in §4; Appendices A, D, and E specify the scoring function, search space, and pseudocode. Appendix B defines the experiment grid and full run configuration. Because coherence and separation serve as both optimization targets and evaluation metrics, the independent LLM-as-judge evaluation (§6.4) is especially important for validating the results. Code will be released upon acceptance. R EFERENCES Radhakrishna Achanta, Appu Shaji, Kevin Smith, Aurelien Lucchi, Pascal Fua, and Sabine Süsstrunk. SLIC superpixels compared to state-of-the-art superpixel methods. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(11):2274–2282, 2012. doi: 10.1109/ TPAMI.2012.120. Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 2623–2631, 2019. doi: 10.1145/3292500.3330701. Dimo Angelov. Top2Vec: Distributed representations of topics. arXiv preprint arXiv:2008.09470, 2020. Jianhong Bai, Zuozhu Liu, Hualiang Wang, Ruizhe Chen, Lianrui Mu, Xiaomeng Li, Joey Tianyi Zhou, Yang Feng, Jian Wu, and Haoji Hu. Towards distribution-agnostic generalized category discovery. In Advances in Neural Information Processing Systems, volume 36, pp. 58625–58647. Curran Associates, Inc., 2023. doi: 10.52202/075280-2555. David M. Blei, Thomas L. Griffiths, Michael I. Jordan, and Joshua B. Tenenbaum. Hierarchical topic models and the nested Chinese restaurant process. In S. Thrun, L. Saul, and B. Schölkopf (eds.), Advances in Neural Information Processing Systems, volume 16. MIT Press, 2003a. David M. Blei, Andrew Y. Ng, and Michael I. Jordan. Latent Dirichlet allocation. Journal of Machine Learning Research, 3:993–1022, 2003b. Ricardo J. G. B. Campello, Davoud Moulavi, and Jörg Sander. Density-based clustering based on hierarchical density estimates. In Advances in Knowledge Discovery and Data Mining (PAKDD 2013), Part II, volume 7819 of Lecture Notes in Artificial Intelligence, pp. 160–172. Springer, 2013. doi: 10.1007/978-3-642-37456-2_14. Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. Efficient intent detection with dual sentence encoders. In Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, pp. 38–45, 2020. doi: 10.18653/v1/2020.nlp4convai-1. 5. 10
Text of page 11
Under review as a conference paper at ICLR 2027 Tianyu Gao, Xingcheng Yao, and Danqi Chen. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6894–6910, Online and Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.552. Maarten Grootendorst. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794, 2022. Zhenqi He, Yuanpei Liu, and Kai Han. SEAL: Semantic-aware hierarchical learning for generalized category discovery. In Advances in Neural Information Processing Systems, volume 38, 2025. URL . Xu Ji, João F. Henriques, and Andrea Vedaldi. Invariant information clustering for unsupervised image classification and segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9865–9874, 2019. Michelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer, and Michael S. Bernstein. Concept induction: Analyzing unstructured text with high-level concepts using LLooM. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24), Honolulu, HI, USA, 2024. Association for Computing Machinery. doi: 10.1145/3613904.3642830. Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. An evaluation dataset for intent classification and out-of-scope prediction. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 1311–1316, 2019. doi: 10.18653/v1/D19-1131. Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281, 2023. Leland McInnes, John Healy, and James Melville. UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. Chau Minh Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, and Mohit Iyyer. TopicGPT: A prompt-based topic modeling framework. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2956– 2984, Mexico City, Mexico, 2024. Association for Computational Linguistics. doi: 10.18653/v1/ 2024.naacl-long.164. Divya Shanmugam, Davis Blalock, Guha Balakrishnan, and John Guttag. Better aggregation in test-time augmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1194–1203, 2021. Evgeny Stepanov, Benoit Favre, Firoj Alam, Shammur Chowdhury, Karan Singla, Jeremy Trione, Frédéric Béchet, and Giuseppe Riccardi. Automatic summarization of call-center conversations. In IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), Demo Papers, Scottsdale, AZ, USA, December 2015. URL . Alex Tamkin, Miles McCain, Kunal Handa, Esin Durmus, Liane Lovitt, Ankur Rathi, Saffron Huang, Alfred Mountfield, Jerry Hong, Stuart Ritchie, Michael Stern, Brian Clarke, Landon Goldberg, Theodore R. Sumers, Jared Mueller, William McEachen, Wes Mitchell, Shan Carter, Jack Clark, Jared Kaplan, and Deep Ganguli. Clio: Privacy-preserving insights into real-world AI use. arXiv preprint arXiv:2412.13678, 2024. Sagar Vaze, Kai Han, Andrea Vedaldi, and Andrew Zisserman. Generalized category discovery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7482–7491, 2022. Vijay Viswanathan, Kiril Gashteovski, Carolin Lawrence, Tongshuang Wu, and Graham Neubig. Large language models enable few-shot clustering. Transactions of the Association for Computational Linguistics, 12:321–333, 2024. doi: 10.1162/tacl_a_00648. 11
Text of page 12
Under review as a conference paper at ICLR 2027 Mengting Wan, Tara Safavi, Sujay Kumar Jauhar, Yujin Kim, Scott Counts, Jennifer Neville, Siddharth Suri, Chirag Shah, Ryen W. White, Longqi Yang, Reid Andersen, Georg Buscher, Dhruv Joshi, and Nagu Rangan. TnT-LLM: Text mining at scale with large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 5836–5847, Barcelona, Spain, 2024. Association for Computing Machinery. doi: 10.1145/3637528.3671647. Zihan Wang, Jingbo Shang, and Ruiqi Zhong. Goal-driven explainable clustering via language descriptions. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 10626–10649, Singapore, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.657. Hanlei Zhang, Hua Xu, Ting-En Lin, and Rui Lyu. Discovering new intents with deep aligned clustering. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14365–14373, 2021. doi: 10.1609/aaai.v35i16.17689. Shun Zhang, Chaoran Yan, Jian Yang, Jiaheng Liu, Ying Mo, Jiaqi Bai, Tongliang Li, and Zhoujun Li. Towards real-world scenario: Imbalanced new intent discovery. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3949–3963, Bangkok, Thailand, 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.217. Yuwei Zhang, Zihan Wang, and Jingbo Shang. ClusterLLM: Large language models as a guide for text clustering. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 13903–13920, Singapore, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.858. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, et al. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Advances in Neural Information Processing Systems, volume 36, 2023. 12
Text of page 13
Under review as a conference paper at ICLR 2027
A
Notation. Let D = {(d i , c i )} N
i=1 be a corpus where each document d i is a single-sentence core
trigger (see §5) pre-assigned to a category c i ∈ C, with |C| ≫ 1 (140+ in our setting). Let ϕ : d ↦→
R m be a pre-trained embedding function (m = 1024). For each category c, define D c = {d i : c i =
c} and the category embedding matrix E c ∈ R |D c |×m .
A.1
I MPLEMENTATION D ETAILS
R ECURSIVE P ROPOSAL P ROCEDURE
For a node v containing n v documents with embeddings E v , the recursive procedure is:
1. Apply UMAP to reduce E v ∈ R n v ×m to Z v ∈ R n v ×d (d = 5, cosine metric)
2. Apply HDBSCAN to Z v , obtaining clusters {C 1 , . . . , C k } and noise set N v
3. If splitting makes no progress, retain the current cluster as a leaf; otherwise recurse on each
C j with |C j | > τ split
4. Each remaining C j with |C j | ≤ τ split becomes a leaf of T c
The minimum cluster size ranges from 20 to 200 in the search, so a 30-document subgroup can only
become a separate proposal if the chosen minimum is at most 30. One UMAP/HDBSCAN configuration is shared across recursion levels. The proposal tree is an intermediate structure; consolidated
definitions and document assignments form the final output.
A.2
I TERATIVE M ERGING P ROTOCOL
We apply greedy merging with iterative convergence:
1. Compute sim δ (l, l ′ ) for all leaf pairs within category c
2. Sort candidates above threshold τ δ by similarity (highest first)
3. Each leaf participates in at most one merge per pass (greedy exclusion)
4. After merging, recompute centroids and delta vectors on the mutated tree
5. Repeat until no new candidates or max iterations reached
Merge candidates are validated by an LLM (sampling 20 sentences from each cluster and querying
for semantic equivalence) in a two-phase protocol: collect all decisions, then batch-mutate. Typical
convergence: 1–3 iterations.
Tree-level Bayesian optimization. Phase 1 will over-fragment or under-fragment depending on
the HDBSCAN parameters. We share a single parameter configuration θ across all depths and use
Optuna’s (Akiba et al., 2019) TPE sampler over 50 trials to find the right degree of fragmentation.
Each trial constructs the complete recursive tree T c (θ) and scores it:
θ ∗ = arg max S(T c (θ))
θ
(2)
The scoring function S is a composite of three terms:
S(T ) = w count · R leaf (T ) − P noise (T ) + B entropy (T )
where R leaf is a Gaussian reward centered on a target leaf-count range [L min , L max ]:
(︃
)︃
(L − µ) 2
L min + L max
L max − L min
R leaf (T ) = exp −
, µ =
, σ =
2
2σ
2
4
(3)
(4)
P noise is an asymmetric noise penalty (over-noise penalized at twice the rate of under-noise), and
B entropy is a normalized entropy bonus that rewards balanced leaf sizes. An optional quality-aware
mode adds topic coherence C and separation S:
C(T ) + S(T )
S quality (T ) = S(T ) + w q ·
(5)
2
13
Text of page 14
Under review as a conference paper at ICLR 2027
When |D c | > 50,000, optimization runs on a random subsample and the final tree is rebuilt on full
data with θ ∗ .
Paraphrase-augmented embeddings. Lexical scatter causes semantically equivalent documents
(“my bill is too high” vs. “I’m being overcharged”) to separate in embedding space, creating spurious
proposals. We tighten representations before recursion by averaging each document’s embedding
with embeddings of K=4 LLM-generated paraphrases (Gao et al., 2021; Shanmugam et al., 2021),
weighted by their cosine similarity to the original (threshold τ s = 0.75; paraphrases below this are
discarded). The augmented embedding is:
ϕ̂(d i ) =
{︃
w k =
cos(ϕ(d i ), ϕ(p ki ))
0
if ≥ τ s
otherwise
(6)
Embeddings use GTE-Large (Li et al., 2023) (1024-dimensional); paraphrases are generated by
Claude Sonnet 4.6.
B
E XPERIMENT G RID AND E NGINEERING D ETAILS
Paraphrase generation. 4 paraphrases per core trigger generated by Claude Sonnet 4.6, precomputed and cached (31,940 total paraphrases for the pilot set).
Experiment grid. We conduct 108 experiments (from 144 combinations, filtered by constraint:
weighted mode requires augmentation enabled):
Table 5: Experiment grid dimensions.
∑︁ K
k
k=1 w k · ϕ(p i )
,
∑︁ K
1 + k=1 w k
ϕ(d i ) +
Dimension
Options
Count
Text representation
Optuna trials
Split threshold
Augmentation
Weighted averaging
Quality optimization
core_trigger, systemic_issue, combined
50, 80
50, 100, 200
enabled, disabled
enabled, disabled
enabled, disabled
3
2
3
2
2
2
Embedding caching. 9 unique embedding configurations are computed once and reused across
all 108 experiments, saving 92% of embedding API calls.
Reported run configuration. Key configuration for the reported 10-category run: no fixed
recursion-depth cap (recursion stops at the size threshold or when splitting makes no progress),
split threshold τ split = 50 (proposed) / 10,000 (flat baseline), 50 tree-level optimization trials per
category, similarity-weighted paraphrase augmentation enabled, quality optimization enabled, minimum leaf size 20, ∼40,000 total samples. The split threshold is a recursion trigger, not a guaranteed
final size bound: any cluster exceeding τ split is recursed into, but the resulting leaves depend on the
density structure found by HDBSCAN at each level. The largest observed leaf (678 documents)
exceeds τ split because HDBSCAN at that recursion level did not find further sub-structure above the
minimum cluster size.
C
A DDITIONAL A NALYSIS
C.1
S AMPLING R ISK
A secondary risk: sampling. Beyond granularity, sampling-based discovery methods face an
additional risk for small subcategories. A subcategory of share p is entirely absent from a uniform
random sample of n documents with probability (1 − p) n . For p = 0.01 and n = 200, this is
approximately 13%. The probability of appearing fewer than three times—too few for a proposer to
14
Text of page 15
Under review as a conference paper at ICLR 2027 recognize a pattern—is approximately 68%. This risk compounds with frequency- and size-based pruning in downstream stages. C.2 W HY D ELTA V ECTORS E NABLE C ONSOLIDATION Within a category, all documents are about roughly the same thing (e.g., billing), so their embeddings cluster tightly around the category centroid µ c . Proposal centroids differ from µ c by small perturbations, and cosine similarity between any two is dominated by the large shared “billing” component. This is the degenerate similarity problem described in §1: without removing the shared direction, consolidation cannot distinguish redundant proposals from genuinely distinct ones. Subtracting the category centroid is analogous to centering data before PCA: it removes the mean and reveals the variance structure underneath. What remains is the direction in which each proposal specializes—the part that distinguishes “unexpected charges” from “autopay failures.” The threshold τ δ controls how aggressively we consolidate. In practice, τ δ ∈ [0.70, 0.90] works across categories: lower values merge more proposals (reducing fragmentation at the risk of collapsing genuinely distinct topics), while higher values are conservative. C.3 S CALABILITY Embeddings are computed once and partitioned by category; each category then runs independently and categories can execute in parallel. For categories exceeding 50,000 samples, Optuna optimization runs on a random subsample and the final tree is rebuilt on full data with the best parameters. Embedding caching across the 108-experiment grid eliminates 92% of embedding API calls. C.4 G RANULARITY AND C OVERAGE Topic counts and coherence scores summarize average quality but do not show how the method changes the concentration of assignments—the size of the largest groups and the fraction of documents left unassigned. Table 6 compares leaf-size ranges across the three methods. Table 6: Leaf-size ranges across 10 categories. Min and max are over all leaves in each method. The proposed method produces smaller maximum groups while avoiding singleton leaves in this configuration; minimum-size constraints (min_cluster_size ≥ 20) partly explain the absence of tiny leaves, and increased rejection partly explains smaller groups. Method Flat LLM-first Proposed (ours) Smallest leaf Largest leaf Topics Unassigned 20 1 22 3,633 1,325 678 111 203 234 15.3% 0.2% 28.6% Where large groups shrink most. The strongest structural signal comes from categories where flat clustering produces very few, very large groups. Table 7 isolates four such categories. Table 7: Topic counts, largest-leaf sizes, and noise rates for four categories with large flat-clustering groups. These category-level comparisons characterize changes in granularity and coverage; they do not establish document-level correspondence between baseline and proposed leaves. Flat Category Service Not Restored Unable to Access Account Update Payment Methods Port-In Request Proposed Topics Largest Noise Topics Largest Noise 3 2 3 12 3,633 2,971 2,750 1,434 1.0% 2.2% 0.5% 2.0% 23 18 29 30 234 516 376 193 37.3% 36.4% 21.3% 23.0% 15
Text of page 16
Under review as a conference paper at ICLR 2027 Service Not Restored illustrates the tradeoff most clearly: the single mega-cluster of 3,633 documents splits into 23 topics, but the unassigned fraction rises from 1% to 37%. The LLM judge rates the finer topics more actionable (mean rises from 3.67 to 4.02), and twice as many topics cross the actionable threshold—but that gain comes at the cost of leaving a third of documents unassigned. Counterexamples. The proposed method does not always increase fragmentation. In three categories it produces fewer topics than flat clustering: Charges & Fees (24 → 21), Multi-Symptom (26 → 22), and Update Account Details (20 → 13), all with lower noise. This suggests the method’s effects depend on category structure rather than simply increasing fragmentation everywhere. The tradeoff is not uniformly favorable: Update Account Details falls from 0.819 to 0.807 coherence and 0.172 to 0.136 separation under the proposed method. Categories where flat clustering already finds many groups may not benefit from recursive splitting. Size of what was found. Median leaf sizes are comparable across methods (Table 6); the structural difference is at the extremes. The largest flat leaf is over five times the largest proposed leaf—the method compresses the biggest groups rather than uniformly shrinking all topics. Per-category topic counts are also more uniform: the max/min ratio drops from 13× (flat) to 2.3× (proposed). In 8 of 10 categories, at least one proposed subcategory holds less than 2% of that category’s volume—the hidden issues the method is designed to surface. C.5 P ER -C ATEGORY B REAKDOWNS Table 8: Per-category breakdown: proposed pipeline vs. LLM-first baseline (10 categories). LLM-first Baseline Category Proposed (ours) Topics Coh. Sep. Topics Coh. Sep. Charges & Fees Activate a Device Make A Payment Multi-Symptom Port-In Request Service Not Restored Unable to Access Account Unable to Call Update Account Details Update Payment Methods 15 10 18 34 29 15 15 12 14 41 0.749 0.760 0.752 0.746 0.777 0.774 0.752 0.751 0.715 0.772 0.080 0.052 0.112 0.129 0.091 0.079 0.088 0.080 0.157 0.115 21 27 24 22 30 23 18 27 13 29 0.805 0.840 0.839 0.820 0.848 0.853 0.847 0.829 0.807 0.863 0.113 0.102 0.137 0.107 0.104 0.109 0.123 0.115 0.136 0.112 Overall 203 0.758 0.098 234 0.835 0.116 D H YPERPARAMETER D ETAILS E A LGORITHM P SEUDOCODE B UILD T REE recursively applies UMAP and HDBSCAN to generate proposals. Oversized clusters are split again until they meet τ split or a splitting attempt makes no progress. In the latter case, the current cluster remains a leaf even if it exceeds the threshold; an unchanged cluster is not passed back into recursion. There is no fixed maximum depth. Noise remains outside the proposal leaves, and all documents are considered during final definition-based assignment (§4.3). 16
Text of page 17
Under review as a conference paper at ICLR 2027 Table 9: Per-category mean actionability scores (1–5 scale, category macro-average) and fraction of topics rated actionable (≥ 4). Categories sorted by the actionability gap between proposed and flat baseline. Flat Proposed Category n Act %≥4 n Act %≥4 n Act %≥4 ∆ Act Unable to Access Account Activate a Device Update Payment Methods Service Not Restored Make A Payment Port-In Request 2 4 3 3 7 12 3.00 3.00 3.83 3.67 3.71 3.50 0.0 0.0 66.7 33.3 42.9 33.3 13 10 31 10 15 24 3.35 3.40 4.10 3.85 3.93 3.46 30.8 30.0 74.2 60.0 53.3 41.7 18 27 29 23 24 30 3.67 3.65 4.43 4.02 3.98 3.60 44.4 51.9 86.2 69.6 58.3 53.3 +0.67 +0.65 +0.60 +0.36 +0.26 +0.10 Unable to make/receive calls Multi-Symptom Charges & Fees Update Account Details 10 26 24 20 3.75 3.54 3.90 3.88 70.0 38.5 62.5 60.0 10 21 15 11 3.50 3.43 3.53 3.27 30.0 23.8 33.3 9.1 27 22 21 13 3.67 3.43 3.69 3.65 44.4 27.3 47.6 46.2 −0.08 −0.11 −0.21 −0.22 Table 10: Search space for tree-level Bayesian optimization. LLM-first Parameter Range Type (UMAP) (UMAP) (HDBSCAN) (HDBSCAN) [10, 50] [0.0, 0.3] [20, 200] [5, 50] Integer Float Integer Integer F P IPELINE A BLATIONS F.1 S CORING F UNCTION C OMPONENTS Each scoring component addresses a specific failure mode. Without the noise penalty, the optimizer finds parameter configurations that push too many documents into the noise set. Without the entropy bonus, it produces lopsided trees where one leaf absorbs most documents. The three terms together steer the optimizer toward trees that are fragmented enough to expose finer distinctions but not so fragmented that the output is unmanageable. F.2 M ERGING C RITERION Raw cosine similarity between centroids within a category is uniformly high (> 0.85) and cannot distinguish redundant proposals from distinct ones. Using raw cosine instead of delta-vector similarity for consolidation would either merge too aggressively (collapsing distinct subcategories) or too conservatively (leaving redundant proposals). Delta-vector projection removes the shared category direction, making genuine redundancy visible. F.3 E MBEDDING S TRATEGY Table 13 isolates the effect of paraphrase augmentation on proposal quality. All three rows use the same pipeline configuration (split threshold, Optuna budget, quality optimization); only the embedding strategy differs. Without augmentation, lexical scatter causes semantically equivalent documents (e.g., “my bill is too high” vs. “I’m being overcharged”) to separate in embedding space, creating spurious proposals. The result is 265 leaves—31 more than the weighted configuration—with lower coherence (0.804 vs. 0.835). Unweighted averaging of paraphrase embeddings reduces fragmentation to 250 leaves and raises coherence to 0.830, confirming that tighter representations before recursion produce fewer spurious splits. Similarity-weighted averaging (Appendix A), which discards paraphrases below a cosine threshold of 0.75 and weights the rest by similarity to the original, further reduces leaves to 17
Text of page 18
Under review as a conference paper at ICLR 2027 Table 11: Default scoring function parameters. Parameter Value Target leaf range [L min , L max ] Target noise band [η min , η max ] Cluster count reward weight w count Noise penalty weight w n Balance bonus weight w b Quality weight w q (optional) [5, 50] [0.05, 0.20] 1.0 1.0 0.2 0.3 Algorithm 1 Clusters-as-Proposals: Recursive Taxonomy Construction Require: Embeddings E c , split threshold τ split , Optuna budget T 1: Initialize Optuna study with TPE sampler 2: for trial t = 1, . . . , T do 3: Sample θ t = (n neighbors , min_dist, min_cluster_size, min_samples) 4: T t ← B UILD T REE (E c , θ t , τ split ) 5: Compute S(T t ) via Equations 3–5 6: Report S(T t ) to Optuna 7: end for 8: θ ∗ ← best parameters from study 9: T ∗ ← B UILD T REE (E c , θ ∗ , τ split ) 10: Apply delta-vector merging (§4.2) to T ∗ 11: return T ∗ 234 and achieves the highest coherence (0.835) and separation (0.116). Separation is comparable across augmentation strategies (0.112 → 0.108 → 0.116), indicating that similarity-weighted averaging primarily benefits coherence and noise reduction rather than inter-cluster distinguishability. G LLM- AS -J UDGE P ROTOCOL For each topic with ≥10 assigned documents, we sample 10 random documents and prompt the judge (Claude Sonnet 4.6, temperature 0.1) to rate the topic on two 1–5 scales: specificity (1 = too broad to act on, 5 = single well-defined problem assignable to one team) and actionability (1 = no clear action possible, 5 = direct, measurable intervention possible). Each topic is evaluated twice with different document samples; per-topic scores are averaged over both runs. Scores are then averaged within each category, and we report the unweighted macro-average across categories as the primary statistic. Topic-level averages are reported separately. Topics with fewer than 10 documents were removed from analysis: all 111 flat and all 234 proposed topics are eligible; 160 of 203 LLM-first topics qualify (the remaining 43 have <10 documents). Significance is assessed with two-sided Wilcoxon signed-rank tests paired by category (n=10). Total evaluation: ∼1,000 LLM calls across all three approaches. G.1 S IZE –A CTIONABILITY D ETAILS The proposed approach shows a significant negative correlation between topic size and actionability (Spearman ρ=−0.215, p=0.001, n=234): smaller topics produced by recursive splitting tend to receive higher actionability ratings. Flat clustering shows a weaker trend (ρ=−0.171, p=0.073), while LLM-first shows no significant correlation (ρ=−0.053, p=0.508). Median leaf sizes are 90 (proposed), 95 (flat), and 129 (LLM-first) over the evaluated topics; the largest difference in concentration is at the extremes (Table 6), not in typical leaf size. G.2 A CTIONABLE T OPIC C OUNTS Of the 234 proposed topics, 127 (54.3%) are rated actionable (≥ 4), compared to 54 of 111 (48.6%) for flat clustering and 68 of 160 (42.5%) for LLM-first. The proposed approach also produces 18
Text of page 19
Under review as a conference paper at ICLR 2027 Table 12: Ablation of tree-level scoring components. Scoring Variant Leaf Target Hit Noise in Band Balance High High High Low High High Low Low High Leaf reward only + Noise penalty + Entropy bonus (full) Table 13: Effect of paraphrase augmentation on the full pipeline (10 categories, same configuration). All reported main-text results use similarity-weighted augmentation (bottom row). Embedding Strategy No augmentation Augmented (unweighted) Augmented (similarity-weighted) Leaves Noise % Coherence Separation 265 250 234 22.0% 21.3% 28.6% 0.804 0.830 0.835 0.112 0.108 0.116 the most high-scoring topics: 51 rated 5 on actionability, versus 14 for flat and 20 for LLM-first. However, these counts should be read as “topics rated actionable,” not as 127 distinct actionable discoveries: the pairwise redundancy evaluation found substantial overlap among nearest-neighbor topics, and several highly rated topics may describe overlapping interventions. Actionability and distinctness must be assessed together. 19
Text of page 20
Under review as a conference paper at ICLR 2027 H A DDITIONAL UMAP V ISUALIZATIONS Figure 3: UMAP projections for three categories comparing flat baseline (left) with recursive proposal generation (right). Flat clustering produces few, large clusters (e.g., 4 clusters for “Activate a Device” with 0% noise), while recursive splitting surfaces fine-grained proposals (40 clusters, 38% noise). Color indicates cluster membership; gray points are noise. 20