---
type: post
title: Clarify a call once, then choose what each reader sees
date: '2026-10-09'
description: We rewrite each customer call once into short, tagged statements. An
  encoder gains from clearer text alone; weaker prompted readers also gain from selection.
authors:
- Mikhail L Arbuzov
- Sisong Bei
paper: https://telegrapher.ai/research/clarify-then-focus.md
html: https://telegrapher.ai/blog/clarify-then-focus/
---

# Clarify a call once, then choose what each reader sees

An agent offers a phone line for $15.00 a month. The customer would rather wait: they are moving, they want the TV service set back up afterwards, and they will get a cell phone then. Later an analyst wants to know which brands came up, what price was offered, why the offer was put off, and what the customer plans next. Each answer sits in a different handful of turns, inside a call that runs to thousands of tokens.

Put questions like these to millions of calls and the reading becomes the bill. A prompted model handed the whole transcript has to find the relevant turns again for every question, wading through disfluencies, pronouns set up minutes earlier, and claims the speech recognizer split across turns. A supervised classifier is cheaper, but it must learn its decision from the same mess. In [the paper](/research/clarify-then-focus/) we pay for the interpretation once per call, then decide for each question how much of it the reader sees.

## A call rewritten into statements a rule can find

Three raw recognizer turns from a synthetic call in the paper:

```
[5]  C: I wanted to add the, uh, the sports package, for the college football this weekend.
[10] A: sure, so it's uh $13.00 a month, but I can do 50% off, so that's $6.50.
[11] C: OK, yeah, that, that sounds great, let's do it.
```

The same turns after rewriting, tags at the end of each line:

```
5.1  C  I want to add the sports package for the college football games this weekend.    request  SPORTS_PACKAGE  actual
10.1 A  The [MULTI_SPORT_PACK] is $13.00 a month, but I can add it at 50% off, which is $6.50.    offer  SPORTS_PACKAGE  actual
11.1 C  Please add the [MULTI_SPORT_PACK].    accept  SPORTS_PACKAGE  actual
```

The filler is gone. Prices come through exactly, and the agent's offer resolves to a canonical product token. "Let's do it" reads as accepting that product because turn 10 had already made the offer: the contract is to resolve references from the turns at hand and invent nothing past them. Every line keeps its speaker and source turn, then carries three tags: what the utterance does, which business topic it concerns, and how the claim is held (an actual event, for instance, or a future intent).

That is statement normalization. Its three tag vocabularies are fixed in advance, with 10 speech acts, 33 subjects and 6 qualifications. It runs before any question is asked, and its output is kept.

Rewriting before tagging is deliberate. In the paper's running example, one customer turn (92) holds a deferral, the move itself and a request to set the TV back up. A single tag on that turn would have to describe all three. Normalized, turn 92 becomes three statements (92.1 to 92.3), shown here with their neighbors and with the place names elided:

```
89.1 A  We are giving you a $15.00 phone line.                   offer    WIRELESS    actual
92.1 C  I will wait until I move before taking the phone line.    refuse   WIRELESS    actual
92.2 C  I am moving from […] to […].                              inform   MOVE        actual
92.3 C  You will have to come set the TV back up after I move.    request  TECH_VISIT  actual
93.1 C  After I move, I will get a cell phone.                    accept   WIRELESS    intent
```

Now a question can pick its evidence with a predicate. `obj = MOVE OR obj = TECH_VISIT` returns 92.2 and 92.3; the customer's stance on the phone line, `speaker = C AND obj = WIRELESS AND (act = refuse OR act = accept)`, returns 92.1 and 93.1. Nothing here calls a model: it is string matching over stored lines. A careless predicate loses evidence just as cheaply. Filter for refusals alone and the future acceptance in 93.1 drops out.

## Clearer text alone helps the encoder

The task is STOP, a call-level label for whether further offers of the pitched product to this customer should be suppressed. No complaint or refusal decides it by itself. Because the label encodes a business policy, it has to come from the people who own that policy, and their time is scarce. Every human score below comes from 66 calls where two annotators reached consensus, 17 of them positive.

Same encoder architecture, two inputs. Raw transcripts gave 0.7879 STOP F1, normalized statements 0.8235, and ROC-AUC rose too. Nothing was selected. A larger run on 16,050 held-out calls points the same way, with the precision gain growing as the recall setting rises. Its labels, though, come from the large teacher model that also wrote the normalized text; that run measures agreement with the teacher, not independent truth.

## Narrowing helps weak readers and hurts some strong ones

Nine prompted readers each got three inputs under the same rubric: the raw transcript, the full normalized call, and the statements whose subject tag matches the product on offer. Four of them, by STOP F1:

| Reader | Raw | Full normalized | Selected |
|---|---|---|---|
| glm-4.7-flash | 0.4545 | 0.5833 | 0.7857 |
| nova-micro | 0.4516 | 0.5405 | 0.7692 |
| claude-haiku-4.5 | 0.8235 | 0.8571 | 0.8235 |
| qwen3-32b | 0.7000 | 0.6250 | 0.5614 |

For the two weakest readers in the table, both steps helped, and selection helped more. Across all nine, the median paired change was +0.0336 from raw to full and +0.1064 from full to selected. Haiku is the counterexample: best on the full call, with narrowing handing the gain back. Qwen got worse at each step.

Clarity and narrowing are separate interventions, then, and which one a reader wants is an empirical question about that reader. With 17 positives, one call moves recall by about 0.059, so small per-reader differences settle nothing.

## A 0.6B normalizer in place of the teacher

The teacher, Llama 4 Maverick, wrote normalized statements and task labels for 194,042 calls. We distilled its normalization into 0.6B and 1.7B students, three seeds per size; seeds varied more than sizes did. Then came the deployment test. The production classifier, trained once on teacher output, had its input switched to the 0.6B student's, with no retraining.

Its F1 slipped from 0.8485 to 0.8387. Precision went up and recall went down, and the interval on the change covers both zero and a loss large enough to matter. The high-recall end is where it shows: at the 0.90 recall setting, precision falls from 0.8421 with teacher input to 0.5484 with student input. Quality has to be checked at the operating point a deployment actually uses.

Still, 0.8387 is above the best input configuration of either Llama 3.3 70B or Maverick as a prompted reader, and a little below the 0.8571 Haiku reached on the full teacher-normalized call. No large model runs anywhere in that serving path.

## What twenty questions would cost

Cost splits into a part paid once per call, for normalization and its metadata, and a part paid for each question. In dollars per million calls, on the paper's batch and list prices:

| Configuration | Once per call | Per question | Total, twenty questions |
|---|---|---|---|
| 0.6B normalizer + encoder | $264.85 | $6.68 | $398 |
| Haiku reading the full normalized call | not counted | $1,915.95 | $38,319 |

At twenty separately scored questions per call that is a 96.2-fold gap; with one question it is about seven-fold. The Haiku row leaves out the teacher normalization its input needs, and counting it would widen the gap.

## What changes for a call-analytics pipeline

The expensive reading moves upstream, to a small model that runs once per call. Each later question adds its own selector and trained head over what is stored, served by a small encoder pass. Tagging already shares work this way in the paper: three tag heads sit on one encoder trunk over the same stored statements.

The paper names our Telegraph English work as the closest earlier proposal. It rewrites text into atomic fact lines, each addressable on its own. The units here are conversational, and shortening is not the goal; what this paper adds is a measurement of selection over those units as an intervention in its own right.

Two of our conversation-analytics papers sit nearby. [High-Load Budgeted Categorization of Customer Care Calls](/research/high-load-call-categorization/) comes at per-call cost from another side: a fine-tuned encoder answers the calls it can and escalates the rest to an LLM. [The Same-Family Halo](/research/same-family-halo/) shows that agreement among LLM labelers can reflect shared labeling preferences rather than independent confirmation. That bears on our silver-label run, where one teacher writes both the input and the labels; the human set is the independent check.

## What is still open

The human evidence is one business policy, in one English customer-service domain, on 66 calls, and consensus filtering may have dropped difficult ones. Nor have we taken the coupled pipeline apart. Without an evidence-matched selection of original turns, surface cleanup, a summary or raw-text retrieval as baselines, we cannot say how much of the gain belongs to rewriting, canonicalization, tagging or selection. Normalization can lose a qualification and selection can drop needed evidence; neither rate has been measured on its own.

The cost comparison has limits of its own. It leaves out labeling, training, calibration and maintenance. The savings over many questions per call are projected; quality over several questions at once has not been measured. We scored one question.

The larger open question concerns the reader. For now a head trained for each question is the stronger route, and each new head brings its own labeling and calibration. A reusable reader, one that takes a task definition in place of task training, could cut that work. We ran the released Laya typed-decisions checkpoint on selected statements: across four wordings of the definition its ROC-AUC ranged from 0.607 to 0.733, against 0.9616 for the task-trained encoder on the full normalized call. Whether a reader of this kind can recover a policy like STOP from its definition alone is still open.
