Calls rewritten once into tagged statements let a 0.6B pipeline near Haiku: 0.8387 vs 0.8571 F1
Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Mikhail L Arbuzov, Sisong Bei, Dmitry Dimov, Karan Dave, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Navita Jain
ICLR 2027 submission, September 2026
Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path.
What we did and found
Normalization rewrites each bounded window of turns into short, self-contained, first-person statements under a fixed contract. Filler and greetings go. A turn carrying several claims becomes several statements; prices and names are restated exactly; known brand variants map to canonical tokens; and every statement keeps the index of the turn it came from. Each statement then gets three tags from fixed vocabularies: a speech act (10 values), a business subject (33), and a qualification such as an actual event or a future intent (6). Construction sees no downstream question, and its output is stored. For a given question, a Boolean predicate over tags, speaker and entity matches can pick out a subset, or a reader can take the whole normalized call. The task is STOP, a call-level policy label for whether further offers of the pitched product to a customer should be suppressed, and three comparisons hang on it. A supervised ModernBERT encoder reads raw or normalized input. Nine prompted readers each see raw, full normalized and selected input under one rubric. And a 0.6B student normalizer, distilled from the teacher, replaces teacher output under a production classifier that is not retrained.
Normalization alone lifted the encoder on the 66 human-consensus calls: STOP F1 went from 0.7879 to 0.8235 and ROC-AUC from 0.9304 to 0.9616, with no selection involved. Scored against teacher labels on 16,050 held-out calls, the same comparison shows precision gains of 9.8 to 15.8 percentage points that grow as the recall setting rises. Prompted readers split. The median paired F1 change was +0.0336 from raw to full normalized input and +0.1064 from full to selected, and for several weaker readers both steps paid: glm-4.7-flash went from 0.4545 to 0.5833 to 0.7857. Haiku, by contrast, did best on the full normalized call, and qwen3-32b lost ground at each step. Swapping in the 0.6B student moved the production classifier from 0.8485 to 0.8387 F1, trading recall for precision; at the 0.90 recall setting, precision fell from 0.8421 to 0.5484. On the reported rates the student pipeline comes to about $398 per million calls at twenty questions per call, against $38,319 for Haiku's reader alone on full teacher-normalized input, the configuration that scored 0.8571.
Key numbers
| Encoder STOP F1, normalized inputModernBERT classifier on normalized statements without selection, against 0.7879 on raw transcripts; 66 human-consensus calls, 17 positives | 0.8235 |
| Median F1 change from selectionpaired change from full to selected normalized input across nine prompted readers; raw to full normalized is +0.0336 | +0.1064 |
| Small-model pipeline STOP F10.6B student normalizer under the unchanged production classifier; teacher input gives 0.8485, Haiku on full teacher-normalized input 0.8571 | 0.8387 |
| Precision at high recall, student inputat the 0.90 recall setting, against 0.8421 with teacher input to the same classifier | 0.5484 |
| Projected serving cost, twenty questionsper million calls for the student pipeline including normalization, against $38,319 for Haiku's reader alone; a 96.2-fold difference on reported batch and list prices | $398 |
What this does not show
The human evidence is one business policy in one English customer-service domain, scored on 66 consensus calls with 17 positives; one positive call moves recall by about 0.059, and consensus filtering may leave out difficult calls. The larger 16,050-call comparison measures agreement with the teacher that also writes the normalized input, not independent correctness. Rewriting, canonicalization, tagging and selection are coupled by design, and separating their contributions would take baselines the paper does not run: an evidence-matched selection of original turns, surface cleanup, a summary, and raw-text retrieval. Source fidelity and tag correctness were not evaluated independently, so it is open how often normalization drops a qualification or selection drops needed evidence. The 0.6B student trails the teacher materially at high recall, and its F1 interval includes practically relevant loss. The costs are conditional estimates on offline batch and list prices. They leave out teacher normalization for the Haiku configuration, along with labeling, training, calibration and maintenance. Savings over several questions per call are projected at the stated rates, and quality over several questions at once has not been measured. A zero-shot reusable decision reader reached ROC-AUC of 0.607 to 0.733, against 0.9616 for the task-trained encoder.
The authors' abstract
Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.
Figure

More in Conversation analytics
Cite
@misc{arbuzov2026clarify,
title = {Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale},
author = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dave, Karan and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Jain, Navita},
year = {2026},
note = {ICLR 2027 submission},
url = {https://telegrapher.ai/research/clarify-then-focus/}
}Builds on
- Choi et al. (2021). Decontextualization: Making sentences stand-alone.
- Chen et al. (2024). Dense X retrieval: What retrieval granularity should we use?
- Bunt et al. (2017). Dialogue act annotation with the ISO 24617-2 standard.
- Kim and Rush (2016). Sequence-level knowledge distillation.
- Arbuzov et al. (2026). Telegraph English: Semantic prompt compression via structured symbolic rewriting.