{
  "slug": "high-load-call-categorization",
  "title": "High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution",
  "short": "HLBC",
  "line": "applied",
  "line_name": "Conversation analytics",
  "part": null,
  "status": "ICLR 2027 submission",
  "date": "2026-09-17",
  "authors": [
    "Mikhail L Arbuzov",
    "Sisong Bei",
    "Dmitry Dimov",
    "Evgeniya Dontsova",
    "Yaodong Hu",
    "Vincent Lao",
    "Karan Dave",
    "Navita Jain"
  ],
  "abstract": "Industry operators route tens of millions of customer-call summaries a year into 100+ fine-grained, long-tailed categories, and at that volume the choice of model per call is itself a cost decision: cheap fine-tuned encoders are unreliable on ambiguous calls, while a strong LLM resolves them but costs one to two orders of magnitude more. We present a deployed hybrid encoder–LLM cascade. A confidence gate on a fine-tuned encoder answers the calls it can and escalates only the rest to the LLM, handing it just the handful of labels the encoder could not separate rather than the full taxonomy. We formulate the two coupled decisions— whether to escalate, and how much of the label space to expose—as one calibrated policy, of which the deployed fixed-gate, fixed-shortlist system is the special case we measure. The economics are the central result. Because the encoder clears roughly 87% of calls on its own, the cascade cuts total cost by more than 90% against running the LLM over the full label list on every call, and shortlisting the escalated call labels cuts their tokens again—the margin that makes the pipeline viable at this volume rather than prohibitive. Shortlisting also carries a prompt-design lesson: the optimized shortlisted prompt is not the full-list prompt with fewer options but a different prompt, and one written for the full taxonomy transfers poorly when reused unchanged—scoring below the full list on human-gold routed calls even when the true label is present. Tuned for the shortlisted regime, grounding each candidate in a synthesized definition lifts conditional accuracy from 42.7% to 63.5%. Past the prompt, the last lever is shortlist recall—set by ranker quality, not by shortlist sizing or further prompt tuning.",
  "tldr": "A confidence gate lets an encoder answer roughly 87% of customer calls; with a five-label shortlist for the rest, cost falls more than 90% below an LLM reading the full list on every call. The shortlist needs its own prompt.",
  "pages": 17,
  "html": "https://telegrapher.ai/research/high-load-call-categorization/",
  "md": "https://telegrapher.ai/research/high-load-call-categorization.md",
  "reader": "https://telegrapher.ai/research/high-load-call-categorization/read/",
  "pdf": "https://telegrapher.ai/papers/high-load-call-categorization/high-load-call-categorization.pdf",
  "arxiv": null,
  "openreview": "https://openreview.net/forum?id=RyiHFBwOvf",
  "post": "https://telegrapher.ai/blog/high-load-call-categorization/",
  "bibtex": "@misc{arbuzov2026highload,\n  title         = {High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution},\n  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Dave, Karan and Jain, Navita},\n  year          = {2026},\n  note          = {ICLR 2027 submission},\n  url           = {https://telegrapher.ai/research/high-load-call-categorization/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "A gated encoder-LLM cascade cuts call-categorization cost by over 90%; the shortlist needs its own prompt"
    },
    {
      "label": "TL;DR",
      "text": "A confidence gate lets an encoder answer roughly 87% of customer calls; with a five-label shortlist for the rest, cost falls more than 90% below an LLM reading the full list on every call. The shortlist needs its own prompt."
    },
    {
      "label": "Method",
      "text": "A fine-tuned DeBERTa-large encoder scores each call summary over 100+ categories and a fixed top-two-margin gate escalates low-confidence calls, with a top-5 shortlist, to an LLM picker (claude-sonnet-4-6, temperature 0); five leakage-free shortlist prompts are compared with the full-list prompt on a human-gold routed slice (n ≈ 96), scored only where the true label is on the shortlist, and cost is reported relative to sending every call to the picker with the full label list."
    },
    {
      "label": "Key result",
      "text": "Calls the encoder gate accepts: 87%; Cost reduction of the cascade: more than 90%; Full-list prompt reused on the shortlist: 42.7%"
    },
    {
      "label": "Why it matters",
      "text": "For teams that classify at volume with an LLM, the paper puts the money and the risk in different places."
    },
    {
      "label": "Limits",
      "text": "The human-gold hard slice is small (n ≈ 96, about ±10 points), so the prompt comparisons are directional, and several dynamic-k gains sit within sampling noise."
    },
    {
      "label": "Status",
      "text": "ICLR 2027 submission, September 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/high-load-call-categorization/read/, PDF /papers/high-load-call-categorization/high-load-call-categorization.pdf"
    }
  ],
  "note": {
    "slug": "high-load-call-categorization",
    "claim_title": "A gated encoder-LLM cascade cuts call-categorization cost by over 90%; the shortlist needs its own prompt",
    "meta_description": "A fine-tuned encoder answers roughly 87% of customer calls; an LLM picks from five labels for the rest. The full-list prompt, reused there, loses accuracy.",
    "tldr": "A confidence gate lets an encoder answer roughly 87% of customer calls; with a five-label shortlist for the rest, cost falls more than 90% below an LLM reading the full list on every call. The shortlist needs its own prompt.",
    "gist": "An operator files tens of millions of customer-call summaries a year under 100+ fine-grained categories, with frequencies spanning several orders of magnitude. A fine-tuned encoder is cheap but unreliable on ambiguous calls; a strong LLM handles them at one to two orders of magnitude more per call. The deployed cascade lets the encoder answer when the margin between its top two scores is wide, and otherwise hands the LLM just the five labels the encoder ranked highest. Roughly 87% of calls stay with the encoder, and the cascade costs more than 90% less than running the LLM over the full label list on every call. The shortlist carried a lesson. Reused unchanged on five candidates, the prompt written for the full list scored below the full list on human-labeled hard calls, even with the true label present. Rebuilt step by step for the shortlist, the prompt lifted conditional accuracy from 42.7% to 63.5%, peaking once each candidate carried a synthesized definition. Beyond the prompt, shortlist recall, set by ranker quality, is the ceiling. The hard slice is small, so the prompt results are directional.",
    "method": "A fine-tuned DeBERTa-large encoder scores each call summary over 100+ categories and a fixed top-two-margin gate escalates low-confidence calls, with a top-5 shortlist, to an LLM picker (claude-sonnet-4-6, temperature 0); five leakage-free shortlist prompts are compared with the full-list prompt on a human-gold routed slice (n ≈ 96), scored only where the true label is on the shortlist, and cost is reported relative to sending every call to the picker with the full label list.",
    "summary_html": [
      "An operator files tens of millions of customer-call summaries a year under 100+ fine-grained categories whose frequencies span several orders of magnitude. A fine-tuned DeBERTa-large encoder scores each summary over the whole taxonomy, and a gate reads the margin between its top two scores. If the margin is wide, the encoder's label stands and the LLM is never called. If it is narrow, the call escalates, and the LLM <em>picker</em> sees just the encoder's top five labels. Two decisions are made per call, then: whether to escalate, and how much of the label space to expose. The paper treats them as one budgeted policy that conformal calibration could set from a single coverage level; what was deployed and measured is the special case of a fixed margin threshold with a fixed top-5 shortlist. Evaluation uses about 900 human-labeled calls held out of encoder training. The prompt comparisons run on the low-confidence routed slice of that set (n ≈ 96) and count a call only when its true label is on the shortlist, so a drop reflects the picker rather than recall. Prompt assets are written from the validation split, frozen, and applied unchanged to test.",
      "The gate carries the economics. At the deployed operating point the encoder accepts roughly 87% of calls; the shortlisted prompt then uses about 0.14× the input tokens of the full-list one, and together they bring cost more than 90% below sending every call to the picker with the full list. Shortlisting was not accuracy-neutral. Reused on five candidates, the prompt written for the full taxonomy scored 42.7%, while the same picker reached 52.1% on the full list. Diagnostics on a larger, silver-labeled routed slice point away from label noise and candidate order: the picker still misses nearly a third of unanimously labeled calls, and its accuracy is flat across the first three ranks. Rebuilt for the shortlist through four incremental changes, the prompt peaked at 63.5% once each candidate carried a synthesized one- or two-sentence definition; pairwise exclusion rules added on top gave some of that back. Then recall takes over. On the hard slice the true label reaches the top five about four times in five, so end-to-end accuracy lands near half. Ordering quality matters more than list length: on the full human-gold set, the encoder's own ranking reaches 90% recall at <code>k = 4</code>, where hybrid retrieval needs <code>k = 15</code>. Resizing the list shifts cost and accuracy only mildly in projection, and per-call sizing helps under some rules and hurts under others."
    ],
    "key_numbers": [
      {
        "label": "Calls the encoder gate accepts",
        "value": "87%",
        "context": "roughly, at the deployed operating point; about 13% are escalated to the LLM picker"
      },
      {
        "label": "Cost reduction of the cascade",
        "value": "more than 90%",
        "context": "against sending every call to the picker with the full list of 100+ categories"
      },
      {
        "label": "Full-list prompt reused on the shortlist",
        "value": "42.7%",
        "context": "conditional accuracy on the human-gold routed slice (n ≈ 96), below the 52.1% the same picker reaches on the full list",
        "bad": true
      },
      {
        "label": "Definition-grounded shortlist prompt",
        "value": "63.5%",
        "context": "same calls and metric; a +20.8-point observed gain over the reused prompt, directional at about ±10 points"
      },
      {
        "label": "End-to-end accuracy on the hard slice",
        "value": "≈ 51%",
        "context": "shortlist recall 0.80 × conditional picker accuracy 0.635; the gap to 63.5% is recall the prompt cannot recover",
        "bad": true
      }
    ],
    "editorial_html": [
      "For teams that classify at volume with an LLM, the paper puts the money and the risk in different places. Most of the saving comes from the gate, so calibration effort belongs on the escalation decision. Shortlisting is a second saving, and the assumption behind it (fewer options cost nothing while the true label is present) failed on exactly the hard calls this system routes. A prompt tuned on the full taxonomy should be re-tuned, not reused, once the picker sees five near-neighbors; of the prompts tested, the one that gave each candidate a short definition scored highest. The rest of the budget goes to ranker recall, ahead of per-call list sizing: hard-negative fine-tuning aimed at the handful of confusable categories behind most misses.",
      "The paper reads this as consistent with earlier findings. Prior work reports that LLMs choose better from fewer options, but those comparisons did not hold the true label's presence fixed, so part of their gain is distractor removal. Hold recall fixed and narrow a strong picker to near-twin candidates, and the task the prompt has to handle changes.",
      "The paper belongs to the group's conversation-analytics line. Its encoder learns from silver labels formed by consensus across several frontier LLMs, and the headline numbers stay on human gold partly because a silver vote may share the picker's model family. <em>The Same-Family Halo</em> studies that kind of dependence head-on: agreement among models can reflect shared labeling preferences rather than independent confirmation, and a gold-free audit makes the dependence measurable. <em>Clusters Are Proposals</em> works further upstream, on where fine-grained categories come from, giving each candidate subcategory a written definition and merging candidates when an LLM confirms they describe the same subcategory. Here, the highest-scoring shortlist prompt was the one that gave each candidate a written definition."
    ],
    "limitations": "The human-gold hard slice is small (n ≈ 96, about ±10 points), so the prompt comparisons are directional, and several dynamic-k gains sit within sampling noise. The diagnostics against label noise and ordering run on silver labels, which are optimistic and partly self-referential when a silver vote shares the picker's model family. One operator, one taxonomy and one language are covered, with one LLM family as the picker, and end-to-end accuracy is measured at the deployed operating points rather than over a sweep. The measured system is the fixed-margin, fixed top-5 special case. That adaptive conformal sets beat fixed routing at matched cost is a framework property and a design expectation, not a head-to-head result; the per-band coverage guarantee holds only as far as calibration data stay exchangeable with production traffic; and the full coverage-level sweep is left to future work. The top-3 and top-10 cost and accuracy figures are projections from an input-token cost model, since only the top-5 point was run. Prompt caching, which could in principle recover most full-list accuracy at near-shortlist marginal cost, is untested, and so is a learned router in place of the margin gate. The call data cannot be released.",
    "concept_terms": [
      "encoder-LLM cascade",
      "confidence-gated escalation",
      "candidate shortlisting",
      "conformal prediction",
      "long-tailed classification",
      "definition-grounded prompting"
    ],
    "references": [
      "Chen et al. (2023). FrugalGPT: How to use large language models while reducing cost and improving performance.",
      "Cheng et al. (2024). E-commerce product categorization with an LLM-based dual-expert classification paradigm.",
      "Vishwakarma et al. (2025). Prune ’n predict (CROQ): Optimizing LLM decision-making with conformal prediction.",
      "Lu et al. (2024). Mitigating boundary ambiguity and inherent bias for text classification in the era of LLMs.",
      "Ding et al. (2025). Conformal prediction for long-tailed classification."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/high-load-call-categorization.txt",
    "verification": {
      "claims": [
        {
          "claim": "claim_title / tldr / gist / key_numbers: the cascade (gate plus five-label shortlist) costs more than 90% less than sending every call to the picker with the full label list",
          "verdict": "CONFIRMED",
          "source_quote": "the cascade cuts total cost by more than 90% against running the"
        },
        {
          "claim": "claim_title / tldr (corrected): the >90% figure needs the shortlist stacked on the gate; the gate alone is about 8x at the 87% point (roughly 87.5%), so the earlier wording that the gate alone cuts cost by over 90% was replaced",
          "verdict": "CONFIRMED",
          "source_quote": "on the gate, this brings the deployed cascade to more than 90% below the cost of sending every call"
        },
        {
          "claim": "tldr / gist / summary / key_numbers: the encoder accepts roughly 87% of calls; about 13% escalate",
          "verdict": "CONFIRMED",
          "source_quote": "encoder auto-accepts roughly 87% of calls and escalates only ∼13% to the picker"
        },
        {
          "claim": "editorial: most of the saving comes from the gate",
          "verdict": "CONFIRMED",
          "source_quote": "The dominant saving is the encoder gate"
        },
        {
          "claim": "gist / summary: tens of millions of call summaries a year",
          "verdict": "CONFIRMED",
          "source_quote": "tens of millions of call-transcript summaries per year into"
        },
        {
          "claim": "gist / summary: 100+ fine-grained categories whose frequencies span several orders of magnitude",
          "verdict": "CONFIRMED",
          "source_quote": "100+ fine-grained categories whose frequencies span several orders of magnitude"
        },
        {
          "claim": "gist: a strong LLM costs one to two orders of magnitude more per call; the encoder is unreliable on ambiguous calls",
          "verdict": "CONFIRMED",
          "source_quote": "ambiguity well but is one to two orders of magnitude more expensive per call"
        },
        {
          "claim": "method / summary: the encoder is a fine-tuned DeBERTa-large scoring all 100+ categories",
          "verdict": "CONFIRMED",
          "source_quote": "A fine-tuned encoder (DeBERTa-large) produces a softmax over the 100+ categories"
        },
        {
          "claim": "gist / summary / method: the gate reads the margin between the top two scores",
          "verdict": "CONFIRMED",
          "source_quote": "the top-two margin"
        },
        {
          "claim": "summary: when the encoder's label is accepted the LLM is never called",
          "verdict": "CONFIRMED",
          "source_quote": "its answer stands; the LLM is never invoked"
        },
        {
          "claim": "gist / summary / method: escalated calls go to the picker with the encoder's top five labels",
          "verdict": "CONFIRMED",
          "source_quote": "a top-5 shortlist at a tuned"
        },
        {
          "claim": "method: picker is claude-sonnet-4-6 at temperature 0",
          "verdict": "CONFIRMED",
          "source_quote": "claude-sonnet-4-6, temperature 0"
        },
        {
          "claim": "summary: the two decisions are formulated as one policy that a single coverage level could set",
          "verdict": "CONFIRMED",
          "source_quote": "thus fixes the escalation rate, the shortlist length, and the cost at once"
        },
        {
          "claim": "summary / limitations: the measured system is the fixed-gate, fixed top-5 special case",
          "verdict": "CONFIRMED",
          "source_quote": "the deployed fixed-gate, fixed-shortlist system is the special case"
        },
        {
          "claim": "summary: about 900 human-labeled calls held out of encoder training",
          "verdict": "CONFIRMED",
          "source_quote": "set of ∼900 calls that was held out of encoder training"
        },
        {
          "claim": "summary / method / key_numbers / limitations: hard routed slice n ≈ 96, about ±10 points",
          "verdict": "CONFIRMED",
          "source_quote": "the low-confidence hard slice is n ≈ 96 (± ∼10pp)"
        },
        {
          "claim": "summary / method: scored only where the true label is on the shortlist, so a drop is a picker effect",
          "verdict": "CONFIRMED",
          "source_quote": "Conditioning on presence isolates the picker from shortlist recall"
        },
        {
          "claim": "summary: prompt assets are validation-derived, frozen and applied unchanged to test",
          "verdict": "CONFIRMED",
          "source_quote": "then frozen and applied unchanged to test"
        },
        {
          "claim": "method: five shortlist prompt arms compared with the full-list prompt on human gold",
          "verdict": "CONFIRMED",
          "source_quote": "Five-arm prompt study scored against human gold labels"
        },
        {
          "claim": "method / key_numbers: cost reported relative to sending every call to the picker with the full list",
          "verdict": "CONFIRMED",
          "source_quote": "We report relative cost against the full-LLM-prompt baseline"
        },
        {
          "claim": "summary: the shortlisted prompt uses about 0.14x the input tokens of the full-list prompt",
          "verdict": "CONFIRMED",
          "source_quote": "the shortlisted prompt uses ≈0.14× the input tokens"
        },
        {
          "claim": "summary: shortlisting was not accuracy-neutral",
          "verdict": "CONFIRMED",
          "source_quote": "this saving is accuracy-neutral whenever the"
        },
        {
          "claim": "summary / key_numbers / meta_description: the full-list prompt reused on five candidates scores 42.7% (Table 1 Arm A)",
          "verdict": "CONFIRMED",
          "source_quote": "applied unchanged to the 5-candidate shortlist, scores only 42.7%"
        },
        {
          "claim": "summary / key_numbers: the same picker reaches 52.1% on the full list (Table 1)",
          "verdict": "CONFIRMED",
          "source_quote": "below the 52.1% the same picker reaches on the full list (Table 1)"
        },
        {
          "claim": "summary / limitations: label-noise and ordering diagnostics run on the larger, silver-scored routed slice",
          "verdict": "CONFIRMED",
          "source_quote": "the larger routed low-confidence slice (∼1,030 calls, scored against"
        },
        {
          "claim": "summary: the picker still misses nearly a third of unanimously labeled calls (68% correct)",
          "verdict": "CONFIRMED",
          "source_quote": "the shortlist picker is still wrong on nearly a third (68% correct)"
        },
        {
          "claim": "summary: accuracy is flat across the first three ranks",
          "verdict": "CONFIRMED",
          "source_quote": "conditional accuracy is flat across ranks 1–3 (60.0/59.2/60.2%)"
        },
        {
          "claim": "summary (corrected): the rebuild is four incremental changes (Arms B to E) layered on the reused prompt A",
          "verdict": "CONFIRMED",
          "source_quote": "Each arm is an incremental, leakage-free change"
        },
        {
          "claim": "summary / key_numbers / editorial: the definition-grounded prompt (Arm D, 63.5%) scored highest of the tested shortlist prompts",
          "verdict": "CONFIRMED",
          "source_quote": "produces the strongest conditional accuracy among the tested shortlist prompts"
        },
        {
          "claim": "summary: definitions are synthesized, one or two sentences each",
          "verdict": "CONFIRMED",
          "source_quote": "grounding each shortlisted candidate in a synthesized 1–2 sentence"
        },
        {
          "claim": "summary: pairwise exclusion rules on top of definitions gave some accuracy back (Table 1: E 58.3% < D 63.5%)",
          "verdict": "CONFIRMED",
          "source_quote": "+ targeted pairwise exclusion rules"
        },
        {
          "claim": "gist (corrected) / key_numbers: the shortlist-adapted prompt lifts conditional accuracy from 42.7% to 63.5% (+20.8 points; recomputed 63.5 - 42.7 = 20.8); the lift spans Arms B, C and D, so it is no longer attributed to definitions alone",
          "verdict": "CONFIRMED",
          "source_quote": "full-list-style baseline to the shortlist-specific, definition-grounded prompt raises conditional accuracy from"
        },
        {
          "claim": "key_numbers: +20.8-point observed gain on the same routed calls",
          "verdict": "CONFIRMED",
          "source_quote": "a +20.8-point observed improvement on the same routed calls"
        },
        {
          "claim": "summary / key_numbers: hard-slice recall@5 is 0.80 (about four in five) and end-to-end accuracy is about 51% (recomputed 0.80 x 0.635 = 0.508)",
          "verdict": "CONFIRMED",
          "source_quote": "0.80 × 0.635 ≈ 51%"
        },
        {
          "claim": "key_numbers: the gap between 63.5% and about 51% is recall the prompt cannot recover",
          "verdict": "CONFIRMED",
          "source_quote": "51% end-to-end figure is attributable to shortlist recall"
        },
        {
          "claim": "summary (scope corrected to the full human-gold set): the encoder reaches 90% recall at k = 4; hybrid retrieval needs k = 15",
          "verdict": "CONFIRMED",
          "source_quote": "To reach 90% recall, retrieval needs k = 15 but the encoder needs k = 4"
        },
        {
          "claim": "summary: retrieval baseline is hybrid (RRF over BM25 + dense)",
          "verdict": "CONFIRMED",
          "source_quote": "baseline (RRF over BM25 + dense)"
        },
        {
          "claim": "summary: ordering quality matters more than list length",
          "verdict": "CONFIRMED",
          "source_quote": "Its primary driver is ordering quality"
        },
        {
          "claim": "summary / limitations: resizing the list moves cost and accuracy only mildly; top-3 and top-10 are projections",
          "verdict": "CONFIRMED",
          "source_quote": "are projected from an input-token cost model"
        },
        {
          "claim": "summary: per-call sizing helps under some rules and hurts under others",
          "verdict": "CONFIRMED",
          "source_quote": "some implementations improve noticeably while others fall below it"
        },
        {
          "claim": "gist: past the prompt, shortlist recall set by ranker quality is the ceiling",
          "verdict": "CONFIRMED",
          "source_quote": "is shortlist recall—set by ranker quality"
        },
        {
          "claim": "editorial: the gate is the money, the shortlist the accuracy risk; calibration effort on the escalation decision",
          "verdict": "CONFIRMED",
          "source_quote": "The gate is the money; the shortlist is the accuracy risk."
        },
        {
          "claim": "editorial: the transfer failure sits on the routed hard calls",
          "verdict": "CONFIRMED",
          "source_quote": "on the low-confidence routed slice rather than uniformly observed"
        },
        {
          "claim": "editorial: invest in ranker recall via hard-negative fine-tuning on the few confusable categories behind most misses",
          "verdict": "CONFIRMED",
          "source_quote": "hard-negative fine-tuning targeted at the ∼8–19 confusable categories that dominate"
        },
        {
          "claim": "editorial (reworded): the paper reads its result as consistent with prior fewer-options findings once conditioning is explicit",
          "verdict": "CONFIRMED",
          "source_quote": "These findings are consistent with ours once the conditioning is made explicit."
        },
        {
          "claim": "editorial: prior gains partly reflect distractor removal",
          "verdict": "CONFIRMED",
          "source_quote": "removing distractors that the model would"
        },
        {
          "claim": "editorial: silver labels come from consensus across several frontier LLMs",
          "verdict": "CONFIRMED",
          "source_quote": "formed by majority consensus across several frontier LLMs"
        },
        {
          "claim": "editorial / limitations: a silver vote may share the picker's model family; silver is optimistic and partly self-referential",
          "verdict": "CONFIRMED",
          "source_quote": "partly self-referential"
        },
        {
          "claim": "editorial: The Same-Family Halo (checked against content/abstracts.yaml): model agreement can reflect shared labeling preferences; a gold-free audit makes the dependence measurable",
          "verdict": "CONFIRMED",
          "source_quote": "agreement among models can reflect shared labeling preferences"
        },
        {
          "claim": "editorial: Clusters Are Proposals (checked against content/abstracts.yaml): candidate subcategories get written definitions and are merged when an LLM confirms they are the same",
          "verdict": "CONFIRMED",
          "source_quote": "the LLM as a candidate subcategory with a written definition"
        },
        {
          "claim": "limitations: several dynamic-k gains are within sampling noise",
          "verdict": "CONFIRMED",
          "source_quote": "are within sampling noise"
        },
        {
          "claim": "limitations: one operator, taxonomy and language; one LLM family as picker",
          "verdict": "CONFIRMED",
          "source_quote": "The study covers a single operator, taxonomy, and language, and one"
        },
        {
          "claim": "limitations: adaptive conformal beating fixed routing is a framework property, not a head-to-head result",
          "verdict": "CONFIRMED",
          "source_quote": "are stated as framework properties and design-time expectations"
        },
        {
          "claim": "limitations: the per-band guarantee is only as good as calibration exchangeability",
          "verdict": "CONFIRMED",
          "source_quote": "the guarantee is only as good as calibration-set exchangeability under production"
        },
        {
          "claim": "limitations: the coverage-level sweep is left to future work",
          "verdict": "CONFIRMED",
          "source_quote": "we leave the full α-sweep to future work"
        },
        {
          "claim": "limitations: prompt caching is untested",
          "verdict": "CONFIRMED",
          "source_quote": "The one untested escape is prompt caching"
        },
        {
          "claim": "limitations: a learned router is not evaluated",
          "verdict": "CONFIRMED",
          "source_quote": "We do not evaluate it here and leave the head-to-head comparison"
        },
        {
          "claim": "limitations: the call data cannot be released",
          "verdict": "CONFIRMED",
          "source_quote": "Data are real customer-call summaries and cannot be released"
        },
        {
          "claim": "references: Chen et al. (2023) FrugalGPT is in the reference list",
          "verdict": "CONFIRMED",
          "source_quote": "FrugalGPT: How to use large language models while"
        },
        {
          "claim": "references: Cheng et al. (2024) dual-expert paradigm is in the reference list",
          "verdict": "CONFIRMED",
          "source_quote": "E-commerce product categorization with an LLM-based dual-expert classification paradigm"
        },
        {
          "claim": "references: Vishwakarma et al. (2025) CROQ is in the reference list",
          "verdict": "CONFIRMED",
          "source_quote": "Optimizing LLM decision-making with conformal prediction"
        },
        {
          "claim": "references: Lu et al. (2024) is in the reference list",
          "verdict": "CONFIRMED",
          "source_quote": "boundary ambiguity and inherent bias for text classification in the era of LLMs"
        },
        {
          "claim": "references: Ding et al. (2025) is in the reference list",
          "verdict": "CONFIRMED",
          "source_quote": "Conformal prediction for long-tailed"
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}