---
type: paper
slug: clarify-then-focus
title: 'Clarify, Then Focus: Statement Normalization for Conversation Analytics at
  Scale'
authors:
- Mikhail L Arbuzov
- Sisong Bei
- Dmitry Dimov
- Karan Dave
- Evgeniya Dontsova
- Yaodong Hu
- Vincent Lao
- Navita Jain
date: '2026-09-17'
status: ICLR 2027 submission
line: Conversation analytics
pages: 21
html: https://telegrapher.ai/research/clarify-then-focus/
pdf: https://telegrapher.ai/papers/clarify-then-focus/clarify-then-focus.pdf
reader: https://telegrapher.ai/research/clarify-then-focus/read/
json: https://telegrapher.ai/api/papers/clarify-then-focus.json
openreview: https://openreview.net/forum?id=gsmr2A3V3f
---

# Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale

## Paper gist

- **Claim:** Calls rewritten once into tagged statements let a 0.6B pipeline near Haiku: 0.8387 vs 0.8571 F1
- **TL;DR:** Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path.
- **Method:** A Llama 4 Maverick teacher normalizes and labels 194,042 English customer-service calls; on a 66-call human-consensus set with 17 positives, the paper compares raw against normalized input for a supervised encoder, raw, full normalized and tag-selected input for nine prompted readers, and teacher against distilled 0.6B student normalization under a fixed production classifier.
- **Key result:** Encoder STOP F1, normalized input: 0.8235; Median F1 change from selection: +0.1064; Small-model pipeline STOP F1: 0.8387
- **Why it matters:** The audience is any team that puts many questions to one large body of calls.
- **Limits:** The human evidence is one business policy in one English customer-service domain, scored on 66 consensus calls with 17 positives; one positive call moves recall by about 0.059, and consensus filtering may leave out difficult calls.
- **Status:** ICLR 2027 submission, September 2026
- **Read:** reader /research/clarify-then-focus/read/, PDF /papers/clarify-then-focus/clarify-then-focus.pdf

## Abstract

Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.

## What we did and found

Normalization rewrites each bounded window of turns into short, self-contained, first-person statements under a fixed contract. Filler and greetings go. A turn carrying several claims becomes several statements; prices and names are restated exactly; known brand variants map to canonical tokens; and every statement keeps the index of the turn it came from. Each statement then gets three tags from fixed vocabularies: a speech act (10 values), a business subject (33), and a qualification such as an actual event or a future intent (6). Construction sees no downstream question, and its output is stored. For a given question, a Boolean predicate over tags, speaker and entity matches can pick out a subset, or a reader can take the whole normalized call. The task is STOP, a call-level policy label for whether further offers of the pitched product to a customer should be suppressed, and three comparisons hang on it. A supervised ModernBERT encoder reads raw or normalized input. Nine prompted readers each see raw, full normalized and selected input under one rubric. And a 0.6B student normalizer, distilled from the teacher, replaces teacher output under a production classifier that is not retrained.

Normalization alone lifted the encoder on the 66 human-consensus calls: STOP F1 went from 0.7879 to 0.8235 and ROC-AUC from 0.9304 to 0.9616, with no selection involved. Scored against teacher labels on 16,050 held-out calls, the same comparison shows precision gains of 9.8 to 15.8 percentage points that grow as the recall setting rises. Prompted readers split. The median paired F1 change was +0.0336 from raw to full normalized input and +0.1064 from full to selected, and for several weaker readers both steps paid: glm-4.7-flash went from 0.4545 to 0.5833 to 0.7857. Haiku, by contrast, did best on the full normalized call, and qwen3-32b lost ground at each step. Swapping in the 0.6B student moved the production classifier from 0.8485 to 0.8387 F1, trading recall for precision; at the 0.90 recall setting, precision fell from 0.8421 to 0.5484. On the reported rates the student pipeline comes to about $398 per million calls at twenty questions per call, against $38,319 for Haiku's reader alone on full teacher-normalized input, the configuration that scored 0.8571.

## Key numbers

| Measure | Value |
|---|---|
| Encoder STOP F1, normalized input | 0.8235 |
| Median F1 change from selection | +0.1064 |
| Small-model pipeline STOP F1 | 0.8387 |
| Precision at high recall, student input | 0.5484 |
| Projected serving cost, twenty questions | $398 |

## Why it matters

The audience is any team that puts many questions to one large body of calls. The paper splits the cost of a question in two. Making the meaning explicit is paid once per call; picking evidence and deciding is paid per question. With the statements stored, serving a new question takes a rule and a small encoder with a head trained for that question, instead of another large-model read of the transcript. Tagging already works this way, with three tag heads sharing one encoder trunk over the same stored statements.

The practical lesson is that readers differ. Clearer text and narrower input are separate interventions: the encoder gains from the first alone, several weak prompted readers gain from both, and Haiku is better off without the second. Which input a reader gets is something to measure for that reader, not a rule to apply.

Telegraph English rewrites text into atomic fact lines and argues that, because each line is independently addressable, the rewrite doubles as a semantic index. Statement normalization carries the addressable-unit idea into conversation without making length the objective, and this paper measures selection over those units as an intervention in its own right. Within the group's conversation-analytics line, The Same-Family Halo shows that agreement among LLM labelers can reflect shared labeling preferences rather than independent confirmation; that bears on the silver-label results here, where one teacher writes both the normalized input and the labels. High-Load Budgeted Categorization of Customer Care Calls comes at per-call cost from another side, with an encoder that answers the calls it can and escalates the rest to an LLM.

## What this does not show

The human evidence is one business policy in one English customer-service domain, scored on 66 consensus calls with 17 positives; one positive call moves recall by about 0.059, and consensus filtering may leave out difficult calls. The larger 16,050-call comparison measures agreement with the teacher that also writes the normalized input, not independent correctness. Rewriting, canonicalization, tagging and selection are coupled by design, and separating their contributions would take baselines the paper does not run: an evidence-matched selection of original turns, surface cleanup, a summary, and raw-text retrieval. Source fidelity and tag correctness were not evaluated independently, so it is open how often normalization drops a qualification or selection drops needed evidence. The 0.6B student trails the teacher materially at high recall, and its F1 interval includes practically relevant loss. The costs are conditional estimates on offline batch and list prices. They leave out teacher normalization for the Haiku configuration, along with labeling, training, calibration and maintenance. Savings over several questions per call are projected at the stated rates, and quality over several questions at once has not been measured. A zero-shot reusable decision reader reached ROC-AUC of 0.607 to 0.733, against 0.9616 for the task-trained encoder.

## Blog post

[Clarify a call once, then choose what each reader sees](https://telegrapher.ai/blog/clarify-then-focus.md)

## Cite

```bibtex
@misc{arbuzov2026clarify,
  title         = {Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale},
  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dave, Karan and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Jain, Navita},
  year          = {2026},
  note          = {ICLR 2027 submission},
  url           = {https://telegrapher.ai/research/clarify-then-focus/}
}
```
