telegrapher

Most calls never reach the LLM, and the ones that do need a different prompt

High-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade SolutionThe paper: summary, reader and PDF

A customer-care operator files tens of millions of call summaries a year, each under one of 100+ intent categories. A few categories, a common payment-related intent among them, carry heavy volume; the rarest sit several orders of magnitude lower. Many have a near neighbor, close enough that human annotators disagree about where one ends and the next begins.

Two kinds of model can do the filing. A fine-tuned encoder costs a fraction of a cent per call and stumbles on the ambiguous ones. A strong instruction-tuned LLM sorts out the ambiguity at one to two orders of magnitude more per call. At this volume, which model sees which call is a budget decision, and running the LLM on everything is out of reach.

In the paper we describe the cascade we deployed to split calls between the two models, and what it taught us about prompting the expensive side.

The encoder already knows which calls are hard

On a clear call, most of the encoder’s probability lands on one label and the runner-up trails far behind. On a hard call, the top two or three labels sit close together. The encoder is not lost there. It is torn between a few.

So we read the gap. Our encoder, a fine-tuned DeBERTa-large, scores each call summary against the whole taxonomy, and a gate checks the margin between its two highest scores. A wide margin means the encoder’s label stands and the LLM is never called. A narrow one escalates the call, and the LLM receives just the five labels the encoder ranked highest. The LLM, which we call the picker, chooses among those five.

That is an encoder–LLM cascade with a shortlist, and it makes two decisions per call: whether to escalate, and how many labels to expose. In the paper we treat the pair as one budgeted policy that conformal calibration could set from a single coverage level. That calibration would also hold a recall floor in each frequency band, so rare categories are not squeezed off the shortlist. What we deployed and measured is the plainest point in that family: one fixed margin threshold and a fixed top-5 shortlist.

The gate is where the money is

At the deployed operating point the encoder accepts roughly 87% of calls by itself. Shortlisting compounds the saving on the rest. The shortlisted prompt needs about a seventh of the full list’s input tokens, but the picker writes a longer reasoning trace against it, so an escalated call comes out just under three times cheaper, not the sevenfold the input tokens alone would suggest. Stack gate and shortlist, and the cascade costs more than 90% less than sending every call to the LLM with the full label list. At tens of millions of calls a year, that margin separates a pipeline that runs from one that cannot ship.

Five labels turned out to be a different question

Shortlisting assumes that if the true label is among the five, fewer options cost nothing. We tested that on the low-confidence calls the gate routes, scored against human labels (n ≈ 96): same picker, same calls, counting a call only when its true label made the shortlist, so any drop belongs to the picker rather than to recall.

The prompt we had written for the full taxonomy, reused unchanged on five candidates, scored 42.7%. Shown all 100+ categories, the same picker scored 52.1%. Fewer options, true answer present, lower accuracy.

Each of those figures carries about ten points of sampling error either way, so the gap is directional. We went after the obvious suspects on a larger routed slice of about 1,030 calls, using LLM-consensus labels where human ones were missing. Neither suspect fully accounts for the gap. On calls the labeling panel agreed on unanimously, the shortlist picker still missed nearly a third, which points away from contested labels; accuracy holds flat across the first three ranks, which points away from candidate order. What remains bunches on a handful of near-twin category pairs.

Earlier work reports the opposite sign: LLMs choose better from fewer options (CROQ; Lu et al.). The two results are compatible. Those comparisons did not hold the true label’s presence fixed, so part of their gain came from removing distractors the model would otherwise have tripped on. Hold presence fixed, and narrowing a strong model to five near-twins changes its job, from scanning a long list to telling neighbors apart. A prompt written for the first job transfers poorly to the second.

Rebuilding the prompt reverses the gap

Starting from that reused prompt, Arm A, we rebuilt the shortlist prompt in four steps, each layered on the last. Every addition was written from the validation split and frozen before test.

Prompt Conditional accuracy
Full list, all 100+ categories (reference) 52.1%
A. Full-list prompt reused on the shortlist, candidate names only 42.7%
B. + calibrated framing and a justified “Other” 51.0%
C. + contrastive reasoning, forced best-to-worst ranking 56.2%
D. + a synthesized definition for each candidate 63.5%
E. + targeted pairwise exclusion rules 58.3%

Calibrated framing and a justified “Other” bring the shortlist nearly level with the full list. Contrastive ranking takes it past. Arm D scored highest: there, each candidate arrives with a one- or two-sentence “typical call” definition, synthesized offline from about 30 validation calls in its category, and conditional accuracy sits 20.8 points above Arm A on the same calls. Exclusion rules stacked on the definitions gave some of that back.

Read the table for direction, not for credit. Each step is smaller than the sampling error on a single figure, and the arms are cumulative, so no arm isolates definitions. On the broader human-labeled set, which easy calls dominate, all five arms land between 74% and 82%; the transfer problem concentrates on the routed slice.

Then recall sets the ceiling

A prompt cannot pick a label that is not on the list. On the hard slice the true label makes the top five about four times in five, so even with Arm D, end-to-end accuracy lands near half. The drop from 63.5% to about half belongs to the ranker.

Ordering matters more than length. On the full human-labeled set, ranking categories by the encoder’s own probabilities reaches 90% recall at 4 candidates, while a hybrid BM25-plus-dense retrieval baseline needs 15. The misses that remain are stubbornly concentrated: half of them fall in 8 categories, mostly confusable near-synonyms.

Resizing the list buys little. Our cost model projects that ten candidates would add 45% to escalated-tier cost for about a point of accuracy. Sizing the list per call from the encoder’s confidence helped under some rules and hurt under others. For why per-call sizing struggles on escalated calls in particular, we have a hypothesis, argued from our data but not proved: a selection effect. The gate already routes on the confidence signal that per-call sizing would read, so by the time a call is escalated, that signal is partly spent.

Where the rest of the budget should go

If you run something like this, the savings come from the gate and the accuracy risk sits with the shortlist. Spend calibration effort on the escalation decision. Treat the shortlisted prompt as a new prompt, tuned for the shortlist, and give the candidates definitions. What is left goes to ranker recall, through hard-negative fine-tuning on the few categories that dominate the misses.

Our encoder learns from silver labels formed by consensus across several frontier LLMs, and one silver vote may share the picker’s model family; that is why the headline numbers sit on human gold. The Same-Family Halo takes that dependence on directly: agreement among models can reflect shared labeling preferences rather than independent confirmation, and a gold-free audit can measure it. In our study, the highest-scoring prompt was the one that added a definition for each candidate. Clusters Are Proposals works further upstream, on where fine-grained categories and their written definitions come from.

What we have not measured yet

The human-labeled hard slice is small, so the prompt comparisons are directional. The checks against label noise and candidate order ran on LLM-consensus labels, which are optimistic. The study covers one operator, one taxonomy, one language, and one LLM family as the picker. The ten-candidate figure is a projection from an input-token model; only the deployed top-5 point was run.

The larger gap is between framework and measurement. We evaluated the fixed-threshold, fixed top-5 special case, so the claim that adaptive conformal sets beat fixed routing at matched cost is a design expectation, not a head-to-head result. The per-band recall guarantee is a property of the construction, good only while calibration data stay exchangeable with production traffic. The sweep over coverage levels, and a learned router in place of the margin gate, are still to run.

The escape we have not tried is prompt caching. A cached full-list prefix could, in principle, show the picker every label again, so recall would no longer cap accuracy, at close to the shortlist’s marginal cost. If caching delivers that, the question moves from how to prompt a short list to whether the list needs to be short.

Read the paperHigh-Load Budgeted Categorization of Customer Care Calls: An Encoder-LLM Cascade SolutionICLR 2027 submission, September 2026