telegrapher

A gated encoder-LLM cascade cuts call-categorization cost by over 90%; the shortlist needs its own prompt

High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution

Mikhail L Arbuzov, Sisong Bei, Dmitry Dimov, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Karan Dave, Navita Jain

ICLR 2027 submission, September 2026

A confidence gate lets an encoder answer roughly 87% of customer calls; with a five-label shortlist for the rest, cost falls more than 90% below an LLM reading the full list on every call. The shortlist needs its own prompt.

What we did and found

An operator files tens of millions of customer-call summaries a year under 100+ fine-grained categories whose frequencies span several orders of magnitude. A fine-tuned DeBERTa-large encoder scores each summary over the whole taxonomy, and a gate reads the margin between its top two scores. If the margin is wide, the encoder's label stands and the LLM is never called. If it is narrow, the call escalates, and the LLM picker sees just the encoder's top five labels. Two decisions are made per call, then: whether to escalate, and how much of the label space to expose. The paper treats them as one budgeted policy that conformal calibration could set from a single coverage level; what was deployed and measured is the special case of a fixed margin threshold with a fixed top-5 shortlist. Evaluation uses about 900 human-labeled calls held out of encoder training. The prompt comparisons run on the low-confidence routed slice of that set (n ≈ 96) and count a call only when its true label is on the shortlist, so a drop reflects the picker rather than recall. Prompt assets are written from the validation split, frozen, and applied unchanged to test.

The gate carries the economics. At the deployed operating point the encoder accepts roughly 87% of calls; the shortlisted prompt then uses about 0.14× the input tokens of the full-list one, and together they bring cost more than 90% below sending every call to the picker with the full list. Shortlisting was not accuracy-neutral. Reused on five candidates, the prompt written for the full taxonomy scored 42.7%, while the same picker reached 52.1% on the full list. Diagnostics on a larger, silver-labeled routed slice point away from label noise and candidate order: the picker still misses nearly a third of unanimously labeled calls, and its accuracy is flat across the first three ranks. Rebuilt for the shortlist through four incremental changes, the prompt peaked at 63.5% once each candidate carried a synthesized one- or two-sentence definition; pairwise exclusion rules added on top gave some of that back. Then recall takes over. On the hard slice the true label reaches the top five about four times in five, so end-to-end accuracy lands near half. Ordering quality matters more than list length: on the full human-gold set, the encoder's own ranking reaches 90% recall at k = 4, where hybrid retrieval needs k = 15. Resizing the list shifts cost and accuracy only mildly in projection, and per-call sizing helps under some rules and hurts under others.

Key numbers

Calls the encoder gate acceptsroughly, at the deployed operating point; about 13% are escalated to the LLM picker87%
Cost reduction of the cascadeagainst sending every call to the picker with the full list of 100+ categoriesmore than 90%
Full-list prompt reused on the shortlistconditional accuracy on the human-gold routed slice (n ≈ 96), below the 52.1% the same picker reaches on the full list42.7%
Definition-grounded shortlist promptsame calls and metric; a +20.8-point observed gain over the reused prompt, directional at about ±10 points63.5%
End-to-end accuracy on the hard sliceshortlist recall 0.80 × conditional picker accuracy 0.635; the gap to 63.5% is recall the prompt cannot recover≈ 51%

What this does not show

The human-gold hard slice is small (n ≈ 96, about ±10 points), so the prompt comparisons are directional, and several dynamic-k gains sit within sampling noise. The diagnostics against label noise and ordering run on silver labels, which are optimistic and partly self-referential when a silver vote shares the picker's model family. One operator, one taxonomy and one language are covered, with one LLM family as the picker, and end-to-end accuracy is measured at the deployed operating points rather than over a sweep. The measured system is the fixed-margin, fixed top-5 special case. That adaptive conformal sets beat fixed routing at matched cost is a framework property and a design expectation, not a head-to-head result; the per-band coverage guarantee holds only as far as calibration data stay exchangeable with production traffic; and the full coverage-level sweep is left to future work. The top-3 and top-10 cost and accuracy figures are projections from an input-token cost model, since only the top-5 point was run. Prompt caching, which could in principle recover most full-list accuracy at near-shortlist marginal cost, is untested, and so is a learned router in place of the margin gate. The call data cannot be released.

Every number above was checked against the paper text.
The authors' abstract

Industry operators route tens of millions of customer-call summaries a year into 100+ fine-grained, long-tailed categories, and at that volume the choice of model per call is itself a cost decision: cheap fine-tuned encoders are unreliable on ambiguous calls, while a strong LLM resolves them but costs one to two orders of magnitude more. We present a deployed hybrid encoder–LLM cascade. A confidence gate on a fine-tuned encoder answers the calls it can and escalates only the rest to the LLM, handing it just the handful of labels the encoder could not separate rather than the full taxonomy. We formulate the two coupled decisions— whether to escalate, and how much of the label space to expose—as one calibrated policy, of which the deployed fixed-gate, fixed-shortlist system is the special case we measure. The economics are the central result. Because the encoder clears roughly 87% of calls on its own, the cascade cuts total cost by more than 90% against running the LLM over the full label list on every call, and shortlisting the escalated call labels cuts their tokens again—the margin that makes the pipeline viable at this volume rather than prohibitive. Shortlisting also carries a prompt-design lesson: the optimized shortlisted prompt is not the full-list prompt with fewer options but a different prompt, and one written for the full taxonomy transfers poorly when reused unchanged—scoring below the full list on human-gold routed calls even when the true label is present. Tuned for the shortlisted regime, grounding each candidate in a synthesized definition lifts conditional accuracy from 42.7% to 63.5%. Past the prompt, the last lever is shortlist recall—set by ranker quality, not by shortlist sizing or further prompt tuning.

On the blogMost calls never reach the LLM, and the ones that do need a different promptA fine-tuned encoder answers roughly 87% of customer calls. The rest go to an LLM with five candidate labels, and that shortlist needs a prompt of its own.

More in Conversation analytics

Cite

@misc{arbuzov2026highload,
  title         = {High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Dave, Karan and Jain, Navita},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/high-load-call-categorization/}
}

Builds on

  1. Chen et al. (2023). FrugalGPT: How to use large language models while reducing cost and improving performance.
  2. Cheng et al. (2024). E-commerce product categorization with an LLM-based dual-expert classification paradigm.
  3. Vishwakarma et al. (2025). Prune ’n predict (CROQ): Optimizing LLM decision-making with conformal prediction.
  4. Lu et al. (2024). Mitigating boundary ambiguity and inherent bias for text classification in the era of LLMs.
  5. Ding et al. (2025). Conformal prediction for long-tailed classification.