{
  "slug": "clarify-then-focus",
  "title": "Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale",
  "short": "CTF",
  "line": "applied",
  "line_name": "Conversation analytics",
  "part": null,
  "status": "ICLR 2027 submission",
  "date": "2026-09-17",
  "authors": [
    "Mikhail L Arbuzov",
    "Sisong Bei",
    "Dmitry Dimov",
    "Karan Dave",
    "Evgeniya Dontsova",
    "Yaodong Hu",
    "Vincent Lao",
    "Navita Jain"
  ],
  "abstract": "Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.",
  "tldr": "Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path.",
  "pages": 21,
  "html": "https://telegrapher.ai/research/clarify-then-focus/",
  "md": "https://telegrapher.ai/research/clarify-then-focus.md",
  "reader": "https://telegrapher.ai/research/clarify-then-focus/read/",
  "pdf": "https://telegrapher.ai/papers/clarify-then-focus/clarify-then-focus.pdf",
  "arxiv": null,
  "openreview": "https://openreview.net/forum?id=gsmr2A3V3f",
  "post": "https://telegrapher.ai/blog/clarify-then-focus/",
  "bibtex": "@misc{arbuzov2026clarify,\n  title         = {Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale},\n  author        = {Arbuzov, Mikhail L and Bei, Sisong and Dimov, Dmitry and Dave, Karan and Dontsova, Evgeniya and Hu, Yaodong and Lao, Vincent and Jain, Navita},\n  year          = {2026},\n  note          = {ICLR 2027 submission},\n  url           = {https://telegrapher.ai/research/clarify-then-focus/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "Calls rewritten once into tagged statements let a 0.6B pipeline near Haiku: 0.8387 vs 0.8571 F1"
    },
    {
      "label": "TL;DR",
      "text": "Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path."
    },
    {
      "label": "Method",
      "text": "A Llama 4 Maverick teacher normalizes and labels 194,042 English customer-service calls; on a 66-call human-consensus set with 17 positives, the paper compares raw against normalized input for a supervised encoder, raw, full normalized and tag-selected input for nine prompted readers, and teacher against distilled 0.6B student normalization under a fixed production classifier."
    },
    {
      "label": "Key result",
      "text": "Encoder STOP F1, normalized input: 0.8235; Median F1 change from selection: +0.1064; Small-model pipeline STOP F1: 0.8387"
    },
    {
      "label": "Why it matters",
      "text": "The audience is any team that puts many questions to one large body of calls."
    },
    {
      "label": "Limits",
      "text": "The human evidence is one business policy in one English customer-service domain, scored on 66 consensus calls with 17 positives; one positive call moves recall by about 0.059, and consensus filtering may leave out difficult calls."
    },
    {
      "label": "Status",
      "text": "ICLR 2027 submission, September 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/clarify-then-focus/read/, PDF /papers/clarify-then-focus/clarify-then-focus.pdf"
    }
  ],
  "note": {
    "slug": "clarify-then-focus",
    "claim_title": "Calls rewritten once into tagged statements let a 0.6B pipeline near Haiku: 0.8387 vs 0.8571 F1",
    "meta_description": "Rewrite each call once into short, tagged statements: an encoder gains with no selection, and tag-based selection also lifts weaker prompted readers.",
    "tldr": "Rewriting each call once into short, attributed, tagged statements improved a supervised encoder without selection, and tag-based selection helped several weaker prompted readers further. With a distilled 0.6B normalizer, no large model sits in the serving path.",
    "gist": "Conversation analytics asks many questions of the same transcripts. A prompted model that reads the whole call for each question redoes the same interpretation every time. Statement normalization does it once per call: the dialogue becomes short, first-person statements that keep their speaker, a reference to the source turn and canonical entity tokens, and each statement is tagged for speech act, subject and qualification. No downstream question is in view while this happens. Later, a Boolean predicate over tags, speaker and entities can select evidence with no further model call. On an offer-suppression label over English customer-service calls, normalization alone raised a ModernBERT classifier from 0.7879 to 0.8235 F1 on a 66-call human-consensus set. Selection helped several weaker prompted readers further; Haiku did best on the full normalized call. A distilled 0.6B normalizer feeding a fixed production classifier reached 0.8387 F1, against 0.8571 for Haiku on full teacher-normalized input, at a projected 96.2-fold lower serving cost for twenty questions per call. All of it rests on one task and a small human set, and the savings over many questions are projected rather than measured.",
    "method": "A Llama 4 Maverick teacher normalizes and labels 194,042 English customer-service calls; on a 66-call human-consensus set with 17 positives, the paper compares raw against normalized input for a supervised encoder, raw, full normalized and tag-selected input for nine prompted readers, and teacher against distilled 0.6B student normalization under a fixed production classifier.",
    "summary_html": [
      "Normalization rewrites each bounded window of turns into short, self-contained, first-person statements under a fixed contract. Filler and greetings go. A turn carrying several claims becomes several statements; prices and names are restated exactly; known brand variants map to canonical tokens; and every statement keeps the index of the turn it came from. Each statement then gets three tags from fixed vocabularies: a speech act (10 values), a business subject (33), and a qualification such as an actual event or a future intent (6). Construction sees no downstream question, and its output is stored. For a given question, a Boolean predicate over tags, speaker and entity matches can pick out a subset, or a reader can take the whole normalized call. The task is <em>STOP</em>, a call-level policy label for whether further offers of the pitched product to a customer should be suppressed, and three comparisons hang on it. A supervised ModernBERT encoder reads raw or normalized input. Nine prompted readers each see raw, full normalized and selected input under one rubric. And a 0.6B student normalizer, distilled from the teacher, replaces teacher output under a production classifier that is not retrained.",
      "Normalization alone lifted the encoder on the 66 human-consensus calls: STOP F1 went from 0.7879 to 0.8235 and ROC-AUC from 0.9304 to 0.9616, with no selection involved. Scored against teacher labels on 16,050 held-out calls, the same comparison shows precision gains of 9.8 to 15.8 percentage points that grow as the recall setting rises. Prompted readers split. The median paired F1 change was +0.0336 from raw to full normalized input and +0.1064 from full to selected, and for several weaker readers both steps paid: glm-4.7-flash went from 0.4545 to 0.5833 to 0.7857. Haiku, by contrast, did best on the full normalized call, and qwen3-32b lost ground at each step. Swapping in the 0.6B student moved the production classifier from 0.8485 to 0.8387 F1, trading recall for precision; at the 0.90 recall setting, precision fell from 0.8421 to 0.5484. On the reported rates the student pipeline comes to about $398 per million calls at twenty questions per call, against $38,319 for Haiku's reader alone on full teacher-normalized input, the configuration that scored 0.8571."
    ],
    "key_numbers": [
      {
        "label": "Encoder STOP F1, normalized input",
        "value": "0.8235",
        "context": "ModernBERT classifier on normalized statements without selection, against 0.7879 on raw transcripts; 66 human-consensus calls, 17 positives"
      },
      {
        "label": "Median F1 change from selection",
        "value": "+0.1064",
        "context": "paired change from full to selected normalized input across nine prompted readers; raw to full normalized is +0.0336"
      },
      {
        "label": "Small-model pipeline STOP F1",
        "value": "0.8387",
        "context": "0.6B student normalizer under the unchanged production classifier; teacher input gives 0.8485, Haiku on full teacher-normalized input 0.8571"
      },
      {
        "label": "Precision at high recall, student input",
        "value": "0.5484",
        "context": "at the 0.90 recall setting, against 0.8421 with teacher input to the same classifier",
        "bad": true
      },
      {
        "label": "Projected serving cost, twenty questions",
        "value": "$398",
        "context": "per million calls for the student pipeline including normalization, against $38,319 for Haiku's reader alone; a 96.2-fold difference on reported batch and list prices"
      }
    ],
    "editorial_html": [
      "The audience is any team that puts many questions to one large body of calls. The paper splits the cost of a question in two. Making the meaning explicit is paid once per call; picking evidence and deciding is paid per question. With the statements stored, serving a new question takes a rule and a small encoder with a head trained for that question, instead of another large-model read of the transcript. Tagging already works this way, with three tag heads sharing one encoder trunk over the same stored statements.",
      "The practical lesson is that readers differ. Clearer text and narrower input are separate interventions: the encoder gains from the first alone, several weak prompted readers gain from both, and Haiku is better off without the second. Which input a reader gets is something to measure for that reader, not a rule to apply.",
      "Telegraph English rewrites text into atomic fact lines and argues that, because each line is independently addressable, the rewrite doubles as a semantic index. Statement normalization carries the addressable-unit idea into conversation without making length the objective, and this paper measures selection over those units as an intervention in its own right. Within the group's conversation-analytics line, <em>The Same-Family Halo</em> shows that agreement among LLM labelers can reflect shared labeling preferences rather than independent confirmation; that bears on the silver-label results here, where one teacher writes both the normalized input and the labels. <em>High-Load Budgeted Categorization of Customer Care Calls</em> comes at per-call cost from another side, with an encoder that answers the calls it can and escalates the rest to an LLM."
    ],
    "limitations": "The human evidence is one business policy in one English customer-service domain, scored on 66 consensus calls with 17 positives; one positive call moves recall by about 0.059, and consensus filtering may leave out difficult calls. The larger 16,050-call comparison measures agreement with the teacher that also writes the normalized input, not independent correctness. Rewriting, canonicalization, tagging and selection are coupled by design, and separating their contributions would take baselines the paper does not run: an evidence-matched selection of original turns, surface cleanup, a summary, and raw-text retrieval. Source fidelity and tag correctness were not evaluated independently, so it is open how often normalization drops a qualification or selection drops needed evidence. The 0.6B student trails the teacher materially at high recall, and its F1 interval includes practically relevant loss. The costs are conditional estimates on offline batch and list prices. They leave out teacher normalization for the Haiku configuration, along with labeling, training, calibration and maintenance. Savings over several questions per call are projected at the stated rates, and quality over several questions at once has not been measured. A zero-shot reusable decision reader reached ROC-AUC of 0.607 to 0.733, against 0.9616 for the task-trained encoder.",
    "concept_terms": [
      "statement normalization",
      "conversation analytics",
      "evidence selection",
      "speech-act tags",
      "sequence-level distillation",
      "offer suppression"
    ],
    "references": [
      "Choi et al. (2021). Decontextualization: Making sentences stand-alone.",
      "Chen et al. (2024). Dense X retrieval: What retrieval granularity should we use?",
      "Bunt et al. (2017). Dialogue act annotation with the ISO 24617-2 standard.",
      "Kim and Rush (2016). Sequence-level knowledge distillation.",
      "Arbuzov et al. (2026). Telegraph English: Semantic prompt compression via structured symbolic rewriting."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/clarify-then-focus.txt",
    "verification": {
      "claims": [
        {
          "claim": "Llama 4 Maverick is the teacher that supplies normalized statements and task labels",
          "verdict": "CONFIRMED",
          "source_quote": "Llama 4 Maverick supplies normalized statements and task labels for the training corpus."
        },
        {
          "claim": "Corpus of 194,042 English customer-service calls with teacher normalization and labels",
          "verdict": "CONFIRMED",
          "source_quote": "The corpus contains 194,042 English customer-service calls with teacher normalization and"
        },
        {
          "claim": "Human evaluation uses a 66-call consensus set with 17 STOP positives",
          "verdict": "CONFIRMED",
          "source_quote": "use the fixed 66-call consensus subset, including 17 STOP positives"
        },
        {
          "claim": "One positive call moves recall by about 0.059",
          "verdict": "CONFIRMED",
          "source_quote": "0.059. Small differences do not establish superiority or equivalence."
        },
        {
          "claim": "Consensus filtering limits the population evaluated and may exclude difficult calls",
          "verdict": "CONFIRMED",
          "source_quote": "filtering limits the population evaluated; one positive call changes recall by approximately"
        },
        {
          "claim": "STOP is a call-level label for whether further offers of the pitched product should be suppressed",
          "verdict": "CONFIRMED",
          "source_quote": "STOP is a call-level operational label indicating whether further offers of"
        },
        {
          "claim": "Normalization rewrites bounded windows of turns into short, self-contained, first-person statements",
          "verdict": "CONFIRMED",
          "source_quote": "bounded window of turns into short, self-contained, first-person statements"
        },
        {
          "claim": "Prices, numbers and names are restated exactly",
          "verdict": "CONFIRMED",
          "source_quote": "Every number, price, date, and name is restated exactly rather than described,"
        },
        {
          "claim": "Every statement keeps the index of its source turn",
          "verdict": "CONFIRMED",
          "source_quote": "Every statement keeps the global index of the turn it came from"
        },
        {
          "claim": "Speech-act tag has 10 values",
          "verdict": "CONFIRMED",
          "source_quote": "(act, 10 values)"
        },
        {
          "claim": "Subject tag has 33 values",
          "verdict": "CONFIRMED",
          "source_quote": "(obj, 33 values)"
        },
        {
          "claim": "Qualification tag has 6 values",
          "verdict": "CONFIRMED",
          "source_quote": "(mod, 6 values)"
        },
        {
          "claim": "Construction sees no downstream question",
          "verdict": "CONFIRMED",
          "source_quote": "In both cases, construction receives no downstream question."
        },
        {
          "claim": "Selection by Boolean predicate needs no further model call",
          "verdict": "CONFIRMED",
          "source_quote": "requires no further model call once the representation exists"
        },
        {
          "claim": "The supervised reader is a ModernBERT encoder",
          "verdict": "CONFIRMED",
          "source_quote": "The supervised reader is a ModernBERT"
        },
        {
          "claim": "Encoder ablation compares same-architecture classifiers on raw vs normalized input",
          "verdict": "CONFIRMED",
          "source_quote": "The encoder ablation compares classifiers of the same architecture trained on"
        },
        {
          "claim": "Encoder STOP F1 0.7879 raw vs 0.8235 normalized on 66 human-consensus calls (Table 1)",
          "verdict": "CONFIRMED",
          "source_quote": "Normalized statements 0.8235"
        },
        {
          "claim": "Encoder ROC-AUC rises from 0.9304 to 0.9616 with no selection",
          "verdict": "CONFIRMED",
          "source_quote": "Normalization increases both reported metrics without selecting evidence."
        },
        {
          "claim": "Silver comparison on 16,050 held-out calls: precision gains of 9.8 to 15.8 points, growing with recall setting",
          "verdict": "CONFIRMED",
          "source_quote": "precision increases by 9.8, 11.9, 15.6, and 15.8"
        },
        {
          "claim": "16,050-call comparison measures agreement with the teacher that writes the normalized input",
          "verdict": "CONFIRMED",
          "source_quote": "agreement with the teacher that also constructs the normalized input."
        },
        {
          "claim": "Nine prompted readers each see raw, full normalized and selected input under one rubric",
          "verdict": "CONFIRMED",
          "source_quote": "Nine prompted readers each receive raw, full normalized, and selected inputs"
        },
        {
          "claim": "Median paired F1 change +0.0336 raw to full, +0.1064 full to selected",
          "verdict": "CONFIRMED",
          "source_quote": "The median paired F1 change is +0.0336 from raw to full input and +0.1064 from full to"
        },
        {
          "claim": "Several weaker readers gain from both normalization and selection",
          "verdict": "CONFIRMED",
          "source_quote": "For several lower-scoring readers, both stages help"
        },
        {
          "claim": "glm-4.7-flash goes 0.4545 to 0.5833 to 0.7857",
          "verdict": "CONFIRMED",
          "source_quote": "0.4545 to 0.5833 to 0.7857"
        },
        {
          "claim": "Haiku does best on the full normalized call",
          "verdict": "CONFIRMED",
          "source_quote": "Haiku performs best on the full normalized call"
        },
        {
          "claim": "qwen3-32b loses ground at each step",
          "verdict": "CONFIRMED",
          "source_quote": "degrades across both transformations"
        },
        {
          "claim": "0.6B student replaces teacher under a production classifier that is not retrained",
          "verdict": "CONFIRMED",
          "source_quote": "switched from teacher output to student output without retraining"
        },
        {
          "claim": "Student pipeline STOP F1 0.8387",
          "verdict": "CONFIRMED",
          "source_quote": "The student pipeline reaches 0.8387 F1"
        },
        {
          "claim": "Teacher-input production classifier STOP F1 0.8485 (Table 3)",
          "verdict": "CONFIRMED",
          "source_quote": "0.8235 0.8485"
        },
        {
          "claim": "Student substitution trades recall for precision",
          "verdict": "CONFIRMED",
          "source_quote": "Student substitution increases precision and reduces recall."
        },
        {
          "claim": "F1 interval includes practically relevant loss",
          "verdict": "CONFIRMED",
          "source_quote": "which includes both zero and practically"
        },
        {
          "claim": "At 0.90 recall, precision 0.5484 (student) vs 0.8421 (teacher)",
          "verdict": "CONFIRMED",
          "source_quote": "precision is 0.5484 with student input"
        },
        {
          "claim": "Teacher-input precision at 0.90 recall is 0.8421; gap is material",
          "verdict": "CONFIRMED",
          "source_quote": "versus 0.8421 with teacher input. This high-recall gap remains material"
        },
        {
          "claim": "Haiku on full teacher-normalized input scores 0.8571",
          "verdict": "CONFIRMED",
          "source_quote": "Haiku on full teacher-normalized input achieves 0.8571"
        },
        {
          "claim": "Student pipeline has no large-model inference call",
          "verdict": "CONFIRMED",
          "source_quote": "yields a serving pipeline with no large-model inference call"
        },
        {
          "claim": "Projected $398 per million calls (student, incl. normalization) vs $38,319 (Haiku) at twenty questions",
          "verdict": "CONFIRMED",
          "source_quote": "about $398 per million calls for the student pipeline, including normalization, versus $38,319"
        },
        {
          "claim": "Haiku figure is its reader cost alone",
          "verdict": "CONFIRMED",
          "source_quote": "for Haiku’s reader alone"
        },
        {
          "claim": "96.2-fold difference",
          "verdict": "CONFIRMED",
          "source_quote": "This is a 96.2-fold difference"
        },
        {
          "claim": "Costs are conditional estimates on offline batch/list prices",
          "verdict": "CONFIRMED",
          "source_quote": "These are conditional estimates on the reported offline batch/list-price basis"
        },
        {
          "claim": "Costs exclude labeling, training, calibration and maintenance",
          "verdict": "CONFIRMED",
          "source_quote": "It excludes labeling, training,"
        },
        {
          "claim": "Cost model separates one-time preparation from per-question decisions",
          "verdict": "CONFIRMED",
          "source_quote": "a cost model that separates onetime preparation from recurring decisions"
        },
        {
          "claim": "Serving query path is a rule and a small encoder",
          "verdict": "CONFIRMED",
          "source_quote": "the query path contains only a rule and a small encoder"
        },
        {
          "claim": "Three tag heads share one encoder trunk over stored statements",
          "verdict": "CONFIRMED",
          "source_quote": "three classification heads share one encoder trunk over the same stored statements"
        },
        {
          "claim": "Multi-question quality not measured",
          "verdict": "CONFIRMED",
          "source_quote": "We have not yet measured end-to-end quality for"
        },
        {
          "claim": "Human evidence is one task in one English customer-service domain",
          "verdict": "CONFIRMED",
          "source_quote": "The human pilot covers one task in one English customer-service"
        },
        {
          "claim": "Missing baselines: evidence-matched original turns, surface cleanup, summary, raw-text retrieval",
          "verdict": "CONFIRMED",
          "source_quote": "An evidence-matched original-turn control, surface cleanup,"
        },
        {
          "claim": "Source fidelity and tag correctness not independently evaluated",
          "verdict": "CONFIRMED",
          "source_quote": "and tagging correctness require independent evaluation"
        },
        {
          "claim": "Zero-shot Laya reader ROC-AUC 0.607 to 0.733 vs 0.9616 task-trained encoder",
          "verdict": "CONFIRMED",
          "source_quote": "ROC-AUC ranges from 0.607 to 0.733, compared with 0.9616"
        },
        {
          "claim": "This paper measures selection as its own intervention and does not treat length as the objective",
          "verdict": "CONFIRMED",
          "source_quote": "We measure selection as its own intervention,"
        },
        {
          "claim": "Telegraph English (other paper): atomic fact lines, each independently addressable, so the rewrite is a semantic index [checked against abstracts.yaml]",
          "verdict": "CONFIRMED",
          "source_quote": "is an independently addressable fact, so the compressed representation is simultaneously"
        },
        {
          "claim": "The Same-Family Halo (other paper): LLM-labeler agreement can reflect shared labeling preferences rather than independent confirmation [abstracts.yaml]",
          "verdict": "CONFIRMED",
          "source_quote": "agreement among models can reflect shared labeling preferences"
        },
        {
          "claim": "High-Load Budgeted Categorization (other paper): encoder answers what it can and escalates the rest to an LLM [abstracts.yaml]",
          "verdict": "CONFIRMED",
          "source_quote": "encoder answers the calls it can and escalates only the rest to the LLM"
        },
        {
          "claim": "Reference: Choi et al. 2021, Decontextualization",
          "verdict": "CONFIRMED",
          "source_quote": "Decontextualization: Making sentences stand-alone."
        },
        {
          "claim": "Reference: Chen et al. 2024, Dense X retrieval",
          "verdict": "CONFIRMED",
          "source_quote": "Dense X retrieval: What retrieval granularity should we"
        },
        {
          "claim": "Reference: Bunt et al. 2017, ISO 24617-2",
          "verdict": "CONFIRMED",
          "source_quote": "Dialogue act annotation with the ISO 24617-2 standard."
        },
        {
          "claim": "Reference: Kim and Rush 2016, Sequence-level knowledge distillation",
          "verdict": "CONFIRMED",
          "source_quote": "Sequence-level knowledge distillation."
        },
        {
          "claim": "Reference: Arbuzov et al. 2026, Telegraph English",
          "verdict": "CONFIRMED",
          "source_quote": "Telegraph English: Semantic prompt compression via structured symbolic rewriting."
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}