---
type: paper
slug: high-load-call-categorization
title: 'High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade
  Solution'
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Dmitry Dimov
- Evgeniya Dontsova
- Yaodong Hu
- Vincent Lao
- Karan Dave
- Navita Jain
date: '2026-09-17'
status: ICLR 2027 submission
line: Conversation analytics
pages: 17
html: https://telegrapher.ai/research/high-load-call-categorization/
pdf: https://telegrapher.ai/papers/high-load-call-categorization/high-load-call-categorization.pdf
reader: https://telegrapher.ai/research/high-load-call-categorization/read/
json: https://telegrapher.ai/api/papers/high-load-call-categorization.json
openreview: https://openreview.net/forum?id=RyiHFBwOvf
---

# High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution

## Paper gist

- **Claim:** A gated encoder-LLM cascade cuts call-categorization cost by over 90%; the shortlist needs its own prompt
- **TL;DR:** A confidence gate lets an encoder answer roughly 87% of customer calls; with a five-label shortlist for the rest, cost falls more than 90% below an LLM reading the full list on every call. The shortlist needs its own prompt.
- **Method:** A fine-tuned DeBERTa-large encoder scores each call summary over 100+ categories and a fixed top-two-margin gate escalates low-confidence calls, with a top-5 shortlist, to an LLM picker (claude-sonnet-4-6, temperature 0); five leakage-free shortlist prompts are compared with the full-list prompt on a human-gold routed slice (n ≈ 96), scored only where the true label is on the shortlist, and cost is reported relative to sending every call to the picker with the full label list.
- **Key result:** Calls the encoder gate accepts: 87%; Cost reduction of the cascade: more than 90%; Full-list prompt reused on the shortlist: 42.7%
- **Why it matters:** For teams that classify at volume with an LLM, the paper puts the money and the risk in different places.
- **Limits:** The human-gold hard slice is small (n ≈ 96, about ±10 points), so the prompt comparisons are directional, and several dynamic-k gains sit within sampling noise.
- **Status:** ICLR 2027 submission, September 2026
- **Read:** reader /research/high-load-call-categorization/read/, PDF /papers/high-load-call-categorization/high-load-call-categorization.pdf

## Abstract

Industry operators route tens of millions of customer-call summaries a year into 100+ fine-grained, long-tailed categories, and at that volume the choice of model per call is itself a cost decision: cheap fine-tuned encoders are unreliable on ambiguous calls, while a strong LLM resolves them but costs one to two orders of magnitude more. We present a deployed hybrid encoder–LLM cascade. A confidence gate on a fine-tuned encoder answers the calls it can and escalates only the rest to the LLM, handing it just the handful of labels the encoder could not separate rather than the full taxonomy. We formulate the two coupled decisions— whether to escalate, and how much of the label space to expose—as one calibrated policy, of which the deployed fixed-gate, fixed-shortlist system is the special case we measure. The economics are the central result. Because the encoder clears roughly 87% of calls on its own, the cascade cuts total cost by more than 90% against running the LLM over the full label list on every call, and shortlisting the escalated call labels cuts their tokens again—the margin that makes the pipeline viable at this volume rather than prohibitive. Shortlisting also carries a prompt-design lesson: the optimized shortlisted prompt is not the full-list prompt with fewer options but a different prompt, and one written for the full taxonomy transfers poorly when reused unchanged—scoring below the full list on human-gold routed calls even when the true label is present. Tuned for the shortlisted regime, grounding each candidate in a synthesized definition lifts conditional accuracy from 42.7% to 63.5%. Past the prompt, the last lever is shortlist recall—set by ranker quality, not by shortlist sizing or further prompt tuning.

## What we did and found

An operator files tens of millions of customer-call summaries a year under 100+ fine-grained categories whose frequencies span several orders of magnitude. A fine-tuned DeBERTa-large encoder scores each summary over the whole taxonomy, and a gate reads the margin between its top two scores. If the margin is wide, the encoder's label stands and the LLM is never called. If it is narrow, the call escalates, and the LLM picker sees just the encoder's top five labels. Two decisions are made per call, then: whether to escalate, and how much of the label space to expose. The paper treats them as one budgeted policy that conformal calibration could set from a single coverage level; what was deployed and measured is the special case of a fixed margin threshold with a fixed top-5 shortlist. Evaluation uses about 900 human-labeled calls held out of encoder training. The prompt comparisons run on the low-confidence routed slice of that set (n ≈ 96) and count a call only when its true label is on the shortlist, so a drop reflects the picker rather than recall. Prompt assets are written from the validation split, frozen, and applied unchanged to test.

The gate carries the economics. At the deployed operating point the encoder accepts roughly 87% of calls; the shortlisted prompt then uses about 0.14× the input tokens of the full-list one, and together they bring cost more than 90% below sending every call to the picker with the full list. Shortlisting was not accuracy-neutral. Reused on five candidates, the prompt written for the full taxonomy scored 42.7%, while the same picker reached 52.1% on the full list. Diagnostics on a larger, silver-labeled routed slice point away from label noise and candidate order: the picker still misses nearly a third of unanimously labeled calls, and its accuracy is flat across the first three ranks. Rebuilt for the shortlist through four incremental changes, the prompt peaked at 63.5% once each candidate carried a synthesized one- or two-sentence definition; pairwise exclusion rules added on top gave some of that back. Then recall takes over. On the hard slice the true label reaches the top five about four times in five, so end-to-end accuracy lands near half. Ordering quality matters more than list length: on the full human-gold set, the encoder's own ranking reaches 90% recall at k = 4, where hybrid retrieval needs k = 15. Resizing the list shifts cost and accuracy only mildly in projection, and per-call sizing helps under some rules and hurts under others.

## Key numbers

| Measure | Value |
|---|---|
| Calls the encoder gate accepts | 87% |
| Cost reduction of the cascade | more than 90% |
| Full-list prompt reused on the shortlist | 42.7% |
| Definition-grounded shortlist prompt | 63.5% |
| End-to-end accuracy on the hard slice | ≈ 51% |

## Why it matters

For teams that classify at volume with an LLM, the paper puts the money and the risk in different places. Most of the saving comes from the gate, so calibration effort belongs on the escalation decision. Shortlisting is a second saving, and the assumption behind it (fewer options cost nothing while the true label is present) failed on exactly the hard calls this system routes. A prompt tuned on the full taxonomy should be re-tuned, not reused, once the picker sees five near-neighbors; of the prompts tested, the one that gave each candidate a short definition scored highest. The rest of the budget goes to ranker recall, ahead of per-call list sizing: hard-negative fine-tuning aimed at the handful of confusable categories behind most misses.

The paper reads this as consistent with earlier findings. Prior work reports that LLMs choose better from fewer options, but those comparisons did not hold the true label's presence fixed, so part of their gain is distractor removal. Hold recall fixed and narrow a strong picker to near-twin candidates, and the task the prompt has to handle changes.

The paper belongs to the group's conversation-analytics line. Its encoder learns from silver labels formed by consensus across several frontier LLMs, and the headline numbers stay on human gold partly because a silver vote may share the picker's model family. The Same-Family Halo studies that kind of dependence head-on: agreement among models can reflect shared labeling preferences rather than independent confirmation, and a gold-free audit makes the dependence measurable. Clusters Are Proposals works further upstream, on where fine-grained categories come from, giving each candidate subcategory a written definition and merging candidates when an LLM confirms they describe the same subcategory. Here, the highest-scoring shortlist prompt was the one that gave each candidate a written definition.

## What this does not show

The human-gold hard slice is small (n ≈ 96, about ±10 points), so the prompt comparisons are directional, and several dynamic-k gains sit within sampling noise. The diagnostics against label noise and ordering run on silver labels, which are optimistic and partly self-referential when a silver vote shares the picker's model family. One operator, one taxonomy and one language are covered, with one LLM family as the picker, and end-to-end accuracy is measured at the deployed operating points rather than over a sweep. The measured system is the fixed-margin, fixed top-5 special case. That adaptive conformal sets beat fixed routing at matched cost is a framework property and a design expectation, not a head-to-head result; the per-band coverage guarantee holds only as far as calibration data stay exchangeable with production traffic; and the full coverage-level sweep is left to future work. The top-3 and top-10 cost and accuracy figures are projections from an input-token cost model, since only the top-5 point was run. Prompt caching, which could in principle recover most full-list accuracy at near-shortlist marginal cost, is untested, and so is a learned router in place of the margin gate. The call data cannot be released.

## Blog post

[Most calls never reach the LLM, and the ones that do need a different prompt](https://telegrapher.ai/blog/high-load-call-categorization.md)

## Cite

```bibtex
@misc{arbuzov2026highload,
  title         = {High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade Solution},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Dave, Karan and Jain, Navita},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/high-load-call-categorization/}
}
```
