{
  "slug": "clusters-are-proposals",
  "title": "Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text",
  "short": "CAP",
  "line": "applied",
  "line_name": "Conversation analytics",
  "part": null,
  "status": "ICLR 2027 submission",
  "date": "2026-09-17",
  "authors": [
    "Navita Jain",
    "Mikhail L Arbuzov",
    "Dmitry Dimov",
    "Evgeniya Dontsova",
    "Yaodong Hu",
    "Vincent Lao",
    "Karan Dave",
    "Sisong Bei"
  ],
  "abstract": "Enterprise call-center analytics often begins with broad topic groups, while the finer distinctions needed for actionable analysis remain hidden. Discovering these distinctions is challenging when interactions within each group are semantically similar, category frequencies are highly uneven, and the number of categories is unknown. The challenge is sharpest when meaningful subcategories are rare: an infrequent but important customer issue may represent only 1% of a broad category’s traffic. Recent LLM pipelines discover categories by proposing candidate definitions from a sample of documents and then assigning every document against them. We argue that these pipelines lose hidden subcategories before an LLM ever sees them: proposals are drawn from random samples, and small candidates are later pruned, absorbed into coarse labels, or removed by minimum-size rules. We introduce clusters as proposals, which changes where proposals come from. Each category is recursively over-fragmented into size-bounded clusters, so every dense region, however small, reaches the LLM as a candidate subcategory with a written definition. Candidates are merged only when an LLM confirms, from their definitions and supporting texts, that they describe the same subcategory, and every document is then labeled against the final definitions. The design rests on an asymmetry: a redundant proposal costs one merge, but a missing proposal cannot be recovered later. On 10 high-volume categories of a production telecommunications call-center corpus, cluster proposals yield 234 subcategories, compared with 111 from flat clustering and 203 from LLM-first discovery, with higher within-subcategory coherence than both baselines in all 10 categories. In 8 of those categories, at least one discovered subcategory holds less than 2% of that category’s calls—the hidden issues the method is designed to surface.",
  "tldr": "Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not.",
  "pages": 20,
  "html": "https://telegrapher.ai/research/clusters-are-proposals/",
  "md": "https://telegrapher.ai/research/clusters-are-proposals.md",
  "reader": "https://telegrapher.ai/research/clusters-are-proposals/read/",
  "pdf": "https://telegrapher.ai/papers/clusters-are-proposals/clusters-are-proposals.pdf",
  "arxiv": null,
  "openreview": "https://openreview.net/forum?id=NoOxt4FXen",
  "post": "https://telegrapher.ai/blog/clusters-are-proposals/",
  "bibtex": "@misc{jain2026clusters,\n  title         = {Clusters Are Proposals: Discovering Hidden Subcategories in Pre-Categorized Text},\n  author        = {Jain, Navita and Arbuzov, Mikhail L and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Dave, Karan and Bei, Sisong},\n  year          = {2026},\n  note          = {ICLR 2027 submission},\n  url           = {https://telegrapher.ai/research/clusters-are-proposals/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "Over-splitting call categories finds 234 subcategories where flat clustering finds 111"
    },
    {
      "label": "TL;DR",
      "text": "Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not."
    },
    {
      "label": "Method",
      "text": "On single-sentence, LLM-distilled statements of each caller's primary reason for calling, from the 10 highest-volume categories of a production telecommunications call-center corpus, embedded with GTE-Large, recursive UMAP and HDBSCAN proposals with centroid-subtracted, LLM-confirmed merging and LLM assignment against written definitions are compared with single-pass flat clustering and an LLM-first baseline (Claude Sonnet 4.6) on embedding coherence and separation, unassigned rate, and LLM-judged specificity and actionability."
    },
    {
      "label": "Key result",
      "text": "Subcategories discovered: 234; Redundancy rate: 4.1%; Mean topic coherence: 0.835"
    },
    {
      "label": "Why it matters",
      "text": "The readers this serves already have a taxonomy."
    },
    {
      "label": "Limits",
      "text": "Everything rests on one enterprise corpus with no reference taxonomy, so the results speak to coherence and granularity, not to recall of true subcategories or semantic completeness; public benchmarks with known fine intents (CLINC150, BANKING77) are named as the test but not run."
    },
    {
      "label": "Status",
      "text": "ICLR 2027 submission, September 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/clusters-are-proposals/read/, PDF /papers/clusters-are-proposals/clusters-are-proposals.pdf"
    }
  ],
  "note": {
    "slug": "clusters-are-proposals",
    "claim_title": "Over-splitting call categories finds 234 subcategories where flat clustering finds 111",
    "meta_description": "Over-split each call category, then let an LLM merge confirmed duplicates: 10 of 244 proposals proved redundant and 234 subcategories remained.",
    "tldr": "Recursively over-splitting each call category, then merging only the pairs an LLM confirms are duplicates, surfaced 234 subcategories against 111 from flat clustering. Redundancy was cheap, 10 merges among 244 proposals; coverage was not.",
    "gist": "Call centers sort calls into broad categories, and the finer issues inside them, some rare, stay hidden. LLM discovery pipelines propose categories from a random sample of documents and later prune small or rare candidates, so a subcategory holding 1% of a category can miss the proposer entirely: a 200-document sample leaves it out about 13% of the time. Clusters as proposals moves the proposal step. Each category is split recursively with UMAP and HDBSCAN until each cluster is at or below a size threshold or stops splitting; every resulting cluster becomes a candidate, an LLM confirms or rejects each merge that geometry suggests, and each call is then checked against the final written definitions. On the 10 highest-volume categories of a production telecommunications corpus, the method yields 234 subcategories, against 111 from flat clustering and 203 from LLM-first discovery, with higher mean coherence than both; 10 of 244 proposals turned out redundant. Coverage is the cost. 28.6% of calls stay unassigned, against 15.3% for flat clustering. With no reference taxonomy, recall of true subcategories is not measured.",
    "method": "On single-sentence, LLM-distilled statements of each caller's primary reason for calling, from the 10 highest-volume categories of a production telecommunications call-center corpus, embedded with GTE-Large, recursive UMAP and HDBSCAN proposals with centroid-subtracted, LLM-confirmed merging and LLM assignment against written definitions are compared with single-pass flat clustering and an LLM-first baseline (Claude Sonnet 4.6) on embedding coherence and separation, unassigned rate, and LLM-judged specificity and actionability.",
    "summary_html": [
      "Recent LLM pipelines for topic discovery share a sound template, <em>propose-then-assign</em>: an LLM proposes categories with written definitions from a sample of documents, and every document is then labeled against them. The sample is where things go wrong. A subcategory holding 1% of a category is missing from a random 200-document sample about 13% of the time, and in about 68% of samples it shows up fewer than three times; later stages then prune rare candidates, fold fine distinctions into coarse labels, or drop clusters below a minimum size. Clusters as proposals changes where the candidates come from. UMAP and HDBSCAN split each category, and any cluster larger than the threshold of 50 documents is split again, with no depth cap, until each piece fits or splitting stops making progress. Every leaf reaches the LLM as a candidate. To find duplicates, cluster centroids are compared after the category centroid is subtracted (<em>delta-vector</em> similarity); for each flagged pair, an LLM reads sampled calls from both clusters and confirms or rejects the merge. Survivors get a name and a one-sentence definition. Each call is then assigned against those definitions, or left unassigned when none fits.",
      "Across the 10 highest-volume categories of a production telecommunications corpus, the method finds 234 subcategories. Flat clustering, with the same embeddings and tuning, finds 111; an LLM-first baseline finds 203. Mean coherence is higher than both, and higher than LLM-first in all 10 categories. Over-splitting turned out cheap: geometry flagged 80 merge candidates among 244 proposals, and the LLM confirmed 10, a 4.1% redundancy rate. In 8 of the 10 categories at least one discovered subcategory holds less than 2% of its category's calls. An LLM judge rates the topics more specific and more actionable than LLM-first's (p=0.002). Against flat clustering the picture is conditional: actionability rises by over half a point on average where flat clustering finds 2–4 groups, and falls in the four categories where it already finds many. The cost shows up in coverage, with 28.6% of calls unassigned against 15.3% for flat clustering and 0.2% for LLM-first."
    ],
    "key_numbers": [
      {
        "label": "Subcategories discovered",
        "value": "234",
        "context": "across 10 call categories, against 111 from flat clustering and 203 from LLM-first discovery"
      },
      {
        "label": "Redundancy rate",
        "value": "4.1%",
        "context": "10 LLM-confirmed merges among 244 proposals; centroid-subtracted geometry had flagged 80 candidates"
      },
      {
        "label": "Mean topic coherence",
        "value": "0.835",
        "context": "mean pairwise cosine similarity within a topic, against 0.792 for flat clustering and 0.758 for LLM-first"
      },
      {
        "label": "Mean actionability (LLM judge)",
        "value": "3.78",
        "context": "1–5 scale, category macro-average, against 3.58 for both flat clustering and LLM-first"
      },
      {
        "label": "Calls left unassigned",
        "value": "28.6%",
        "context": "against 15.3% for flat clustering and 0.2% for LLM-first",
        "bad": true
      }
    ],
    "editorial_html": [
      "The readers this serves already have a taxonomy. Their categories route calls and set staffing well enough; the trouble sits inside them. A payment-app regression shipped last week, or eSIM provisioning failing on one device model, may come to thirty calls out of three thousand, and by the time it is frequent enough to notice it has been costing customers and agents for weeks. The paper puts the loss before any LLM reads a call. Each pruning rule in earlier pipelines protects something reasonable, whether label quality, cost or privacy; stacked together, they can remove a small group before the proposer ever sees it. Nothing about the LLM changes here. What changes is what it gets shown.",
      "The paper argues, without testing it, that the design applies wherever documents arrive pre-sorted into categories and finer ones are needed. The rule underneath is simple: if a later stage can delete a redundant candidate but nothing downstream can recover a missing one, generate too many and spend judgment on the merge. Geometry cannot make that judgment alone. Inside a narrow category, raw cosine similarity between cluster centroids stays above 0.85, so duplicates and mere neighbors look alike; subtracting the category centroid leaves the direction in which each cluster specializes, and that is what gets compared. Even so, of 80 flagged candidates the LLM judged 70 to describe different subcategories.",
      "The work sits in the group's conversation-analytics line. <em>High-Load Budgeted Categorization of Customer Care Calls</em> routes customer-call summaries into fine-grained, long-tailed categories and finds that, on the escalated calls the LLM sees with a shortlist of labels, grounding each candidate label in a synthesized definition lifts conditional accuracy; this paper comes at the same problem from the other end, discovering finer subcategories inside broad ones and writing a definition for each. <em>The Same-Family Halo</em> adds a caution that applies directly. Agreement among models can reflect shared labeling preferences rather than independent confirmation, and the judge in this evaluation is the same model that runs the LLM-first baseline."
    ],
    "limitations": "Everything rests on one enterprise corpus with no reference taxonomy, so the results speak to coherence and granularity, not to recall of true subcategories or semantic completeness; public benchmarks with known fine intents (CLINC150, BANKING77) are named as the test but not run. Coherence and separation are computed in the same embedding space the clustering uses, and in the reported run they also enter the tuning objective, which favors the embedding-based methods over LLM-first. The 4.1% redundancy rate counts the merges the LLM accepted, 10 against 244 proposals, under a prompt set to keep borderline distinctions; a separate pairwise check found substantial overlap among nearest-neighbor topics, and the paper warns that counts of highly rated topics overstate distinct actionable discoveries. Coverage differs sharply, 28.6% unassigned against 0.2% for LLM-first, and no comparison is made at matched coverage. The judge, Claude Sonnet 4.6, is also the LLM-first baseline's model, and its ratings measure LLM-judged usefulness rather than business outcomes. Against flat clustering the judged gain is not significant (p=0.16) and reverses in four categories; coherence falls in at least one of them, Update Account Details. The count of 8 categories with a sub-2% subcategory is not reported for the baselines. A minimum cluster size still bounds how small a proposal can be, and recursion does not always isolate one problem: a Make a Payment topic of 197 calls still mixes several causes.",
    "concept_terms": [
      "clusters as proposals",
      "hidden subcategories",
      "recursive over-fragmentation",
      "delta-vector consolidation",
      "propose-then-assign",
      "call-center topic discovery"
    ],
    "references": [
      "Pham et al. (2024). TopicGPT: A prompt-based topic modeling framework.",
      "Wan et al. (2024). TnT-LLM: Text mining at scale with large language models.",
      "Grootendorst (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure.",
      "Campello et al. (2013). Density-based clustering based on hierarchical density estimates.",
      "Tamkin et al. (2024). Clio: Privacy-preserving insights into real-world AI use."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/clusters-are-proposals.txt",
    "verification": {
      "claims": [
        {
          "claim": "Clusters as proposals yields 234 subcategories across the 10 categories",
          "verdict": "CONFIRMED",
          "source_quote": "cluster proposals yield 234 subcategories"
        },
        {
          "claim": "Flat clustering finds 111 subcategories and LLM-first discovery 203",
          "verdict": "CONFIRMED",
          "source_quote": "compared with 111 from flat clustering and 203 from LLM-first discovery"
        },
        {
          "claim": "Evaluation covers the 10 highest-volume categories of a production telecommunications call-center corpus",
          "verdict": "CONFIRMED",
          "source_quote": "Detailed results are reported on the 10 highest-volume categories"
        },
        {
          "claim": "A 1% subcategory is entirely absent from a random 200-document sample about 13% of the time (recomputed: 0.99^200 = 0.134)",
          "verdict": "CONFIRMED",
          "source_quote": "has a 13% chance of being entirely absent from a random sample of 200"
        },
        {
          "claim": "It appears fewer than three times in about 68% of such samples (recomputed binomial: 0.677)",
          "verdict": "CONFIRMED",
          "source_quote": "a 68% chance of appearing fewer than three times"
        },
        {
          "claim": "Later stages prune rarely generated candidates, fold fine distinctions into coarse labels, or drop clusters below a minimum size",
          "verdict": "CONFIRMED",
          "source_quote": "clusters below a minimum size are dropped outright"
        },
        {
          "claim": "Each pruning rule protects label quality, cost or privacy",
          "verdict": "CONFIRMED",
          "source_quote": "it protects label quality, cost, or privacy"
        },
        {
          "claim": "Proposals come from UMAP followed by HDBSCAN, recursively splitting oversized clusters",
          "verdict": "CONFIRMED",
          "source_quote": "applying UMAP followed by HDBSCAN"
        },
        {
          "claim": "Split threshold of 50 documents with no depth cap",
          "verdict": "CONFIRMED",
          "source_quote": "a split threshold of 50 with no fixed recursion-depth cap"
        },
        {
          "claim": "Recursion stops at the size threshold or when splitting makes no progress (gist corrected from 'until its clusters are small'; the largest leaf is 678 documents)",
          "verdict": "CONFIRMED",
          "source_quote": "each branch stops when its cluster meets the size threshold or further splitting makes no progress"
        },
        {
          "claim": "Every resulting leaf reaches the LLM as a candidate",
          "verdict": "CONFIRMED",
          "source_quote": "every resulting cluster is shown to the LLM"
        },
        {
          "claim": "Merge candidates are found by comparing centroids after subtracting the category centroid (delta-vector similarity)",
          "verdict": "CONFIRMED",
          "source_quote": "Subtracting the category centroid removes the shared direction"
        },
        {
          "claim": "For each flagged pair an LLM reads sampled calls from both clusters and confirms or rejects the merge",
          "verdict": "CONFIRMED",
          "source_quote": "reading 20 sampled documents from each cluster"
        },
        {
          "claim": "Surviving proposals get a short name and a one-sentence definition",
          "verdict": "CONFIRMED",
          "source_quote": "the LLM generates a short name and a one-sentence definition"
        },
        {
          "claim": "Each call is assigned against the definitions or left unassigned when none fits",
          "verdict": "CONFIRMED",
          "source_quote": "leaves the document unassigned when none applies"
        },
        {
          "claim": "Flat clustering uses the same embeddings and tuned configuration in a single pass",
          "verdict": "CONFIRMED",
          "source_quote": "This baseline uses the same embeddings and the same Optuna-tuned UMAP/HDBSCAN configuration"
        },
        {
          "claim": "Inputs are single-sentence, LLM-distilled statements of the caller's primary reason (method corrected from 'LLM summaries')",
          "verdict": "CONFIRMED",
          "source_quote": "LLM distills each transcript into a single-sentence core trigger"
        },
        {
          "claim": "Embeddings are GTE-Large",
          "verdict": "CONFIRMED",
          "source_quote": "GTE-Large (Li et al., 2023) (1024-dimensional, L2-normalized)"
        },
        {
          "claim": "The LLM-first baseline uses Claude Sonnet 4.6",
          "verdict": "CONFIRMED",
          "source_quote": "LLM-first discovery: A frontier LLM (Claude Sonnet 4.6)"
        },
        {
          "claim": "The judge is Claude Sonnet 4.6, the same model as the LLM-first baseline",
          "verdict": "CONFIRMED",
          "source_quote": "an independent judge (Claude Sonnet 4.6)"
        },
        {
          "claim": "Coherence is the mean pairwise cosine similarity among documents in a topic",
          "verdict": "CONFIRMED",
          "source_quote": "the mean pairwise cosine similarity among all documents assigned to the same topic"
        },
        {
          "claim": "Mean coherence 0.835 vs 0.792 (flat) and 0.758 (LLM-first): higher than both",
          "verdict": "CONFIRMED",
          "source_quote": "Table 2 Coherence column: Flat 0.792, LLM-first 0.758, Proposed (ours) 0.835"
        },
        {
          "claim": "Coherence higher than LLM-first in all 10 categories (recomputed cell by cell from Table 8)",
          "verdict": "CONFIRMED",
          "source_quote": "Coherence improves over the LLM-first baseline in all 10 categories"
        },
        {
          "claim": "Geometry flagged 80 merge candidates among 244 proposals",
          "verdict": "CONFIRMED",
          "source_quote": "Delta-vector similarity flags 80 of 244 proposals as geometric merge candidates"
        },
        {
          "claim": "The LLM confirmed 10 merges, a 4.1% redundancy rate (recomputed 10/244 = 4.10%)",
          "verdict": "CONFIRMED",
          "source_quote": "the LLM confirms only 10"
        },
        {
          "claim": "The LLM judged the other 70 candidates to describe different subcategories",
          "verdict": "CONFIRMED",
          "source_quote": "The remaining 70 candidates share similar deviation directions"
        },
        {
          "claim": "Raw cosine similarity between centroids within a category stays above 0.85",
          "verdict": "CONFIRMED",
          "source_quote": "Raw cosine similarity between centroids within a category stays above 0.85"
        },
        {
          "claim": "In 8 of 10 categories at least one discovered subcategory holds less than 2% of its category's calls",
          "verdict": "CONFIRMED",
          "source_quote": "at least one discovered subcategory holds less than 2% of that"
        },
        {
          "claim": "The sub-2% count is not reported for the baselines (no baseline figure appears anywhere in the text)",
          "verdict": "CONFIRMED",
          "source_quote": "at least one proposed subcategory holds less than 2%"
        },
        {
          "claim": "The judge rates proposed topics more specific and more actionable than LLM-first's, p=0.002",
          "verdict": "CONFIRMED",
          "source_quote": "consistent enough to reach significance on both measures (two-sided Wilcoxon, p=0.002;"
        },
        {
          "claim": "Mean actionability 3.78 vs 3.58 for both baselines, 1-5 scale, category macro-average (recomputed from Table 9: 3.779 / 3.578 / 3.582)",
          "verdict": "CONFIRMED",
          "source_quote": "Table 4 Actionability column: Flat 3.58 (3.69), LLM-first 3.58 (3.63), Proposed (ours) 3.78 (3.80)"
        },
        {
          "claim": "Where flat clustering finds 2-4 groups, actionability rises by over half a point on average (Table 9: mean of +0.67, +0.65, +0.60, +0.36 = 0.57)",
          "verdict": "CONFIRMED",
          "source_quote": "actionability rises by over half a point on average"
        },
        {
          "claim": "Actionability falls against flat in the four categories where flat already finds many groups",
          "verdict": "CONFIRMED",
          "source_quote": "In the four categories where flat clustering already finds many well-separated groups, the proposed"
        },
        {
          "claim": "Against flat the judged gain is not significant (p=0.16)",
          "verdict": "CONFIRMED",
          "source_quote": "the signal is weaker (p=0.16,"
        },
        {
          "claim": "Unassigned 28.6% vs 15.3% for flat clustering",
          "verdict": "CONFIRMED",
          "source_quote": "at the cost of higher noise (28.6% vs. 15.3%)"
        },
        {
          "claim": "Unassigned 28.6% vs 0.2% for LLM-first",
          "verdict": "CONFIRMED",
          "source_quote": "Noise, however, rises substantially (28.6% vs. 0.2%"
        },
        {
          "claim": "Coherence falls against flat in Update Account Details",
          "verdict": "CONFIRMED",
          "source_quote": "Update Account Details falls from 0.819 to 0.807 coherence"
        },
        {
          "claim": "Single enterprise corpus with no reference taxonomy; recall against ground truth not established",
          "verdict": "CONFIRMED",
          "source_quote": "Our evidence comes from a single enterprise corpus with no reference taxonomy"
        },
        {
          "claim": "CLINC150 and BANKING77 are named as the public test but not run",
          "verdict": "CONFIRMED",
          "source_quote": "the public test explicitly: CLINC150"
        },
        {
          "claim": "Coherence and separation are both tuning targets and evaluation metrics, and favor embedding-aligned methods",
          "verdict": "CONFIRMED",
          "source_quote": "Because coherence and separation serve as both optimization targets and evaluation metrics"
        },
        {
          "claim": "No comparison at matched coverage",
          "verdict": "CONFIRMED",
          "source_quote": "These comparisons do not establish quality at matched coverage"
        },
        {
          "claim": "Judge ratings measure LLM-judged usefulness, not business outcomes",
          "verdict": "CONFIRMED",
          "source_quote": "measure LLM-judged usefulness, not demonstrated business outcomes"
        },
        {
          "claim": "The minimum cluster size bounds how small a proposal can be",
          "verdict": "CONFIRMED",
          "source_quote": "a 30-document subgroup can only become a separate proposal if the chosen minimum is at most 30"
        },
        {
          "claim": "A Make a Payment topic of 197 calls still mixes several causes",
          "verdict": "CONFIRMED",
          "source_quote": "proposed topic 32 (197 documents) still combines outdated cards"
        },
        {
          "claim": "The design is argued, not tested, to apply wherever documents are pre-sorted into categories",
          "verdict": "CONFIRMED",
          "source_quote": "anywhere documents are pre-sorted into categories and finer sub-topics are needed"
        },
        {
          "claim": "Thirty calls out of three thousand example (payment-app regression, eSIM provisioning)",
          "verdict": "CONFIRMED",
          "source_quote": "may account for thirty calls out of three thousand"
        },
        {
          "claim": "The 4.1% rate reflects a merge prompt set to keep borderline distinctions (limitation added from source)",
          "verdict": "CONFIRMED",
          "source_quote": "preserving borderline distinctions rather than collapsing"
        },
        {
          "claim": "A pairwise check found substantial overlap among nearest-neighbor topics; actionable counts overstate distinct discoveries (limitation added from source)",
          "verdict": "CONFIRMED",
          "source_quote": "the pairwise redundancy evaluation found substantial overlap among nearest-neighbor"
        },
        {
          "claim": "High-Load Budgeted Categorization routes customer-call summaries into fine-grained, long-tailed categories (checked against its abstract)",
          "verdict": "CONFIRMED",
          "source_quote": "customer-call summaries a year into 100+ fine-grained, long-tailed categories"
        },
        {
          "claim": "High-Load: grounding each candidate in a synthesized definition lifts conditional accuracy in the shortlisted regime (corrected: 'conditional' and the escalated-call scope added)",
          "verdict": "CONFIRMED",
          "source_quote": "grounding each candidate in a synthesized definition lifts conditional accuracy"
        },
        {
          "claim": "The Same-Family Halo: agreement among models can reflect shared labeling preferences rather than independent confirmation (checked against its abstract)",
          "verdict": "CONFIRMED",
          "source_quote": "agreement among models can reflect shared labeling preferences rather than independent confirmation"
        },
        {
          "claim": "All five references (Pham 2024, Wan 2024, Grootendorst 2022, Campello 2013, Tamkin 2024) appear in the paper's reference list",
          "verdict": "CONFIRMED",
          "source_quote": "Reference list: TopicGPT (Pham et al.), TnT-LLM (Wan et al.), BERTopic (Grootendorst), Density-based clustering (Campello et al.), Clio (Tamkin et al.)"
        },
        {
          "claim": "claim_title (234 vs 111) is entailed by the summary after trimming",
          "verdict": "CONFIRMED",
          "source_quote": "cluster proposals yield 234 subcategories"
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}