telegrapher

Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale

Back to the paper page. ICLR 2027 submission, September 2026.

PDF

All 21 pages are shown below.

Page 1 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 1 of 21
Text of page 1
Under review as a conference paper at ICLR 2027

Clarify, Then Focus: Statement Normalization
for Conversation Analytics at Scale

Anonymous authors
Paper under double-blind review

Abstract

Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and
identifying which information matters, repeating costly interpretive work
across the same transcripts. We propose a simple principle: clarify the text,
then focus the reader. Statement normalization transforms dialogue into
short, speaker-attributed statements with source references and semantic
tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full
representation or a relevant subset, depending on what helps them make
the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker
prompted readers benefit from both normalization and selection. A small
model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across
questions supports an inference pipeline built entirely from small models,
making analytics over millions of conversations substantially less expensive.

1

Introduction

A conversation happens once. Its meaning may be reconstructed many times.

An agent offers a Boost Mobile phone line for $15.00 per month, mentioning its partnership
with Dish. The customer wants to wait until after a move, asks for the TV service to be set
up again, and plans to get a cell phone then. Which brands were discussed? What price was
offered? Why was the offer deferred, and what does the customer intend to do next? Each
question draws on a different part of the same exchange. Before answering, a model must
first find these few turns within a conversation spanning thousands of tokens, then separate
the offer, the customer’s circumstances, and their future plans.

At enterprise scale, this interpretive work is repeated across large collections of long form
transcripts. A direct prompted approach asks a language model to read the entire conversation for each task. A supervised classifier can make predictions more cheaply, but must
learn its decision from transcripts containing disfluencies, implicit references, and evidence
distributed across turns. The input representation therefore affects what the downstream
model must do.

This motivates a different division of work: clarify the conversation once, then focus the
reader for each decision. Normalization decomposes dialogue into short, simple, attributed
statements with source references and canonical entity names. Semantic tags make the
statements selectable. The representation is constructed without a downstream question
and stored for subsequent analysis.

This division of work supports a pipeline built entirely from small models. On our offersuppression pilot, a 0.6B normalizer and a task-trained encoder achieve 0.839 F1, compared
with 0.857 for Haiku on full teacher-normalized input. At the reported rates, the pipeline
has approximately 100-fold lower projected serving cost than Haiku’s reader alone at twenty
separately scored questions per conversation (Section 6).

1

Reviewers: please read the Reviewer Guidelines (iclr.cc/Conferences/2027/ReviewerGuidelines) and the AI Policy for Reviewers (iclr.cc/Conferences/2027/AIPolicyForReviewers).
If you used AI to expand, edit, or polish your review, please provide the input text to the LLM. Better still, consider skipping the LLM and submitting your original text: we,
and the authors, are much more interested in your unedited thoughts than in what an LLM has to say. AI-assisted or not, you are putting your name and reputation behind
your review: LLM-generated falsehoods, hallucinations or misrepresentations are subject to disciplinary action, which may include desk-rejecting all papers you have authored.
Page 2 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 2 of 21
Text of page 2
Under review as a conference paper at ICLR 2027

Figure 1: Clarify once, focus for each question. Normalization and tagging run once per
conversation and are stored. For each question, a Boolean predicate selects statements for a
prompted reader, or a task-trained encoder reads all statements. The panels use an excerpt
of Example 1.

Example 1. One exchange, nine tagged statements. The statements span six source turns
(88–93). A denotes the agent and C the customer. Each line ends with tags in the order
speech act · subject · qualification.

88.1 A You can get a $15.00/month phone line with the [BOOST_MOBILE]
promo. offer WIRELESS actual
88.2 A [BOOST_MOBILE] is a partner of [DISH]; you are a valued [DISH]
customer. inform WIRELESS actual

89.1 A We are giving you a $15.00 phone line. offer WIRELESS actual

90.1 C Is the phone line a home phone or a cell phone? ask WIRELESS actual
91.1 A The phone line can be a home phone or a cell
phone. inform WIRELESS actual
92.1 C I will wait until I move before taking the phone
line. refuse WIRELESS actual

92.2 C I am moving from Wadena to Minnesota. inform MOVE actual
92.3 C You will have to come set the TV back up after I
move. request TECH_VISIT actual

93.1 C After I move, I will get a cell phone. accept WIRELESS intent

These units are “greppable.” A wireless tag retrieves the offer and the customer’s responses
even when they do not name either brand. Combining subject, speaker, and speech act
isolates the agent’s offers; other predicates expose the move, TV setup, or future intent. A
downstream model can read all statements or a selected subset, according to what helps its
decision.

Normalization may appear to require a large model. Its output, however, follows a stable
contract: separate claims, preserve speakers and qualifications, and resolve references into
explicit statements. We treat this as a constrained semantic mapping that a small model can
learn from a teacher. This makes both construction and consumption of the representation
candidates for inexpensive inference.

We evaluate the design on an offer-suppression task in customer-service calls. Our contributions connect the representation to its downstream use:

1. Statement normalization with canonical entities. We develop attributed units intended to capture individual claims, preserving source references and canonical
names for known entities. These units make conversational information easier for
downstream models to read and for rules to address.

2
Page 3 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 3 of 21
Text of page 3
Under review as a conference paper at ICLR 2027

2. Programmatic selection through semantic tags. Speech-act, subject, and qualification tags combine with speaker and entity predicates to expose fine-grained evidence
slices. New combinations can be selected from the stored representation without
another generative pass.

3. Utility for encoder and decoder readers. We compare raw and normalized encoder
inputs and three input representations across nine prompted readers. The encoder
benefits without selection; several weaker prompted readers benefit from both normalization and further selection.

4. Small models for preparation and prediction. We substitute a 0.6B normalizer
under an unchanged production classifier and measure the resulting quality tradeoff. We connect this serving configuration to a cost model that separates onetime preparation from recurring decisions. Three tagging heads demonstrate shared
encoder computation; downstream reuse across questions is projected.

2

Related work

Explicit statements and rewriting. Decontextualization aims to preserve a sentence’s meaning while making it interpretable outside its original context (Choi et al., 2021). Discourseaware simplification separates propositions while retaining their relationships (Niklaus et al.,
2022); MinIE represents polarity, modality, attribution, and quantities alongside compact
relational tuples (Gashteovski et al., 2017). These relationships and qualifications matter when decomposition could detach a claim from its conditions. Dense X Retrieval uses
self-contained propositions as retrieval units and trains their producer from large-model examples (Chen et al., 2024). It establishes precedents for both decomposition and a learned
producer. Our evaluation concerns attributed conversation statements with canonical entities and controlled tags as inputs to call-level decisions.

Rewriting also provides an interface to downstream models. Su et al. (2019) recover omitted and referential content in contextual utterances and report gains in dialogue systems.
This targets a current request; we construct an inventory of statements for subsequent questions. Telegraph English rewrites text into compact, addressable units and evaluates smaller
readers, but its QA experiments do not test its proposed tag-based selection mechanism (Arbuzov et al., 2026). A companion study on multi-hop question answering finds that such
readable symbolic re-expression preserves entity content more densely than matched-budget
deletion, truncation, or a coherent prose summary (Bei et al., 2026). We explicitly compare
the full representation with selected statements and include an encoder that benefits without
selection.

Metadata and selective access. Dialogue-act annotation distinguishes communicative functions and qualifiers on functional segments (Bunt et al., 2017). Metadata can also support
retrieval: A-MEM enriches interactions with descriptions, keywords, and tags, embeds the
records, and retrieves by semantic similarity (Xu et al., 2025). SimpleMem combines selective retention, write-time synthesis, and model-planned retrieval over semantic, lexical, and
symbolic indexes for long-term agent memory (Liu et al., 2026). Our controlled tags serve
a simpler access interface: fixed Boolean predicates over statements, speakers, and entities.
Their value is evaluated through the behavior of the downstream reader.

Compression and reuse. RECOMP uses query-focused compression in its QA setting (Xu
et al., 2024); AttentionRAG selects original sentences using query-conditioned attention
(Fang et al., 2025). Query-independent alternatives include LLMLingua-2’s learned word
selection (Pan et al., 2024) and ReadAgent’s reusable gists with question-dependent passage lookup (Lee et al., 2024). Offline extraction for later QA (Fleischman et al., 2003) and
EVAPORATE’s structured views and generated extraction functions (Arora et al., 2023),
and precomputed retrieval indexes such as GraphRAG and RAPTOR (Edge et al., 2024;
Sarthi et al., 2024), likewise move work ahead of individual questions. Query independence
alone is therefore not our contribution. We investigate the combination of readable statements, controlled selection, encoder and decoder utility, and a small learned producer.

3
Page 4 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 4 of 21
Text of page 4
Under review as a conference paper at ICLR 2027

Reusable decision readers. Jev accepts textual state and question-defined output schemas
to produce typed decisions (Almeida, 2026). Laya provides an open-weight Jev-compatible
interface, with a ModernBERT backbone and a learned decision head (Convai Innovations,
2026a). These readers make it possible to specify new decisions without fitting a separate
head for each question. Our representation supplies their evidence; our Laya experiment
tests transfer to a domain-specific policy. Appendix A develops these comparisons and
additional related work.

3

Method

3.1

Constructing attributed statements

A conversation is hard to read because a single claim is rarely contained in a single turn.
It may be split across turns by the speech recognizer, buried under disfluencies, or carried
by a pronoun whose referent was established minutes earlier. Normalization rewrites each
bounded window of turns into short, self-contained, first-person statements, so that a claim
can be read on its own without replaying the call around it. Filler, backchannel, and
greetings are dropped; a turn that runs several claims together becomes one statement
each, and a sentence the recognizer split across turns is rejoined under the turn where it
begins. Every number, price, date, and name is restated exactly rather than described,
redaction placeholders such as [CREDIT_DEBIT_NUMBER] that arrive already typed
from an upstream pass are copied character for character, and known brand variants are
mapped to canonical tokens such as [BOOST_MOBILE] and [DISH] by a deterministic
gazetteer. Every statement keeps the global index of the turn it came from: in Example 1,
the integer identifies that source turn and the suffix distinguishes statements within it. The
discipline throughout is to resolve and restate, never to invent or expand — a pronoun is
replaced only by an antecedent the window already supplies, a short utterance is not inflated
into an intent the surrounding turns do not establish, and each speaker’s stance is kept, so
a refusal stays a refusal.

A few turns from a synthetic call show the effect; Appendix B.5 works the whole exchange
end to end. Three raw recognizer turns,

• [5] C: I wanted to add the, uh, the sports package, for the college football this
weekend.

• [10] A: sure, so it’s uh $13.00 a month, but I can do 50% off, so that’s $6.50.

• [19] C: no, I’m good for now.

become self-contained statements, shown here with the tags of Example 1:

5.1 C I want to add the sports package for the college football games this
weekend. request SPORTS_PACKAGE actual

10.1 A The [MULTI_SPORT_PACK] is $13.00 a month, but I can add it at
50% off, which is $6.50. offer SPORTS_PACKAGE actual

19.1 C I am not interested in [BOOST_MOBILE] right
now. refuse WIRELESS actual

Filler is dropped, the price is restated exactly, the agent’s brand is canonicalized to
[MULTI_SPORT_PACK], and ”no, I’m good for now” is resolved into a refusal of the
[BOOST_MOBILE] service named a turn earlier — kept as a refusal rather than softened
into a deferral. A reader can act on any one of these statements without replaying the
call around it. Because the same fixed rules apply to every window, normalization is a
constrained mapping a small distilled model can learn (Section 3.3). The normalizer itself
performs no typing, entity extraction, or numbering, which are separate downstream stages.

Each statement then carries three tags drawn from fixed vocabularies. The speech-act tag
(act, 10 values) records what the utterance does, such as offer, refuse, or accept. The
subject tag (obj, 33 values) records the business topic, such as WIRELESS, MOVE, or

4
Page 5 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 5 of 21
Text of page 5
Under review as a conference paper at ICLR 2027

TECH_VISIT. The qualification tag (mod, 6 values) records how the claim is held, for
example as an actual event or a future intent. Together the three families describe up to
10 × 33 × 6 = 1,980 statement types. This counts only the controlled tags; the canonical
entity tokens — nine closed keys such as product ([BOOST_MOBILE]) and competitor
([VERIZON]), each with an open vocabulary of values — add a further selection axis through
the literal-match predicates in Appendix B.1. The tag vocabularies are fixed in advance by
the normalization contract rather than learned from data.

Tags label normalized statements rather than raw turns, and this ordering matters. A raw
turn often mixes several claims: customer turn 92 in Example 1 contains a deferral, a move,
and a service request. A single tag on that turn would have to describe all three at once.
Normalization first separates the claims, so that each tag describes one of them.

Tags can come from two producers. The generator can append them while writing each
statement, or a separate encoder can assign them afterwards, using one shared trunk with
a classification head per family. The two routes differ in cost and in accuracy, so the
chosen producer must be counted in the construction cost and its output checked against
the contract. In both cases, construction receives no downstream question. The schema
is still domain-specific: its subject vocabulary names the recurring topics of this business,
which is what makes later selection possible.

3.2

Selecting evidence and predicting

For transcript X, let Z = N (X) be the normalized representation. A task-specific selector
S q supplies evidence to reader H q :

(
)
ŷ q = H q S q (Z) .

Here ŷ q is the reader’s predicted decision for question q. S q either retains all statements
or applies Boolean predicates over tags, speaker, and canonical-entity matches. Selection
requires no further model call once the representation exists. In Example 1, speaker = A
AND act = offer AND obj = WIRELESS retrieves the two offers (88.1 and 89.1). A stance
decision may need both the deferral in 92.1 and future acceptance in 93.1; selecting only
refusals would omit relevant evidence. Appendix B.1 gives exact slices and discusses context
dependencies.

Two kinds of reader consume the representation. The supervised reader is a ModernBERT
encoder with a task-specific classification head (Warner et al., 2025), trained on task labels.
The prompted readers are decoder models given a fixed task rubric and no task-specific
training. Either kind can be fed the full normalized call or the subset chosen by a subjectmatch selector — the STOP selector for the offer-suppression task of Section 4, which keeps
statements whose subject tag matches the offered product; which input serves a given reader
is decided by task performance, not by a fixed rule. In the ablations below, the encoder is
compared on raw and full normalized input, and each prompted reader additionally receives
the selected statements.

Comparing inputs within each reader separates the effect of normalization from the effect
of selection. Selection is an input variant a reader uses when it helps, not a prerequisite for
benefiting from normalization. Section 5 reports which readers gain from it.

3.3

Learning a small producer

Llama 4 Maverick supplies normalized statements and task labels for the training corpus.
The statements supervise smaller normalizers through teacher-output training (Kim & Rush,
2016); the task labels separately supervise classifiers. The normalizer learns the representation, while the classifier learns a decision over it.

We distill two student normalizers, of 0.6B and 1.7B parameters, and train each size with
three seeds. Several seeds are needed because, at this corpus size, a single run cannot
separate a real difference between sizes from training randomness.

5
Page 6 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 6 of 21
Text of page 6
Under review as a conference paper at ICLR 2027

The serving comparison asks the deployment question directly. The production classifier
is trained once on teacher-normalized text and then held fixed. At test time, its input is
switched from teacher output to student output without retraining. This mirrors practice,
where the normalizer may be replaced without retraining every downstream classifier. One
rendering detail is controlled. The teacher emits per-statement turn numbers during normalization and the student does not, so both inputs receive identically synthesized numbering.
Re-rendering the teacher input this way reproduces its original score within a pre-declared
tolerance, so a remaining gap reflects the normalizer rather than formatting.

4

Experimental setup

Task and data. STOP is a call-level operational label indicating whether further offers of
the pitched product to this customer should be suppressed. A complaint, product mention,
or refusal tag alone does not define the decision. Evaluation measures agreement with that
label.

The corpus contains 194,042 English customer-service calls with teacher normalization and
teacher task labels. The production classifier is trained on 151,914 calls. The raw-versus-normalized ablation uses 16,050 held-out calls that fit the context limit in both representations, covering 96.6% of distinct held-out call IDs. These silver-label results measure
agreement with the teacher that also constructs the normalized input.

Independent human evaluation starts from 80 calls annotated by two people. All reported
human scores use the fixed 66-call consensus subset, including 17 STOP positives. Consensus
filtering limits the population evaluated; one positive call changes recall by approximately
0.059. Small differences do not establish superiority or equivalence.

The human set is small by necessity, and this reflects the setting. The STOP label encodes a
business policy, namely when the company should stop offering a product. It must therefore
come from domain experts who own that policy, and general crowd annotators cannot supply
it. Expert time is scarce, so enterprise deployments typically combine a large model-labeled
corpus with a small expert-labeled set for validation. We design the analysis around that
regime.

Comparisons. The encoder ablation compares classifiers of the same architecture trained on
raw and normalized inputs. Its normalized-input classifier is distinct from the production
classifier used for normalizer substitution. Nine prompted readers each receive raw, full normalized, and selected inputs under the same task rubric and denominator. This separates
raw-to-full and full-to-selected changes within each reader. Finally, the unchanged production classifier receives teacher or student normalization. The separate exploratory zero-shot
protocol is reported in Appendix D.

5

Results

The results follow the proposed decomposition: clarify the input for an encoder, focus it for
prompted readers, and replace the large producer with a small one. We then test whether a
reusable decision reader can apply the policy without task-specific training. Human scores
share the 66-call denominator; the larger silver-label analysis is identified separately.

5.1

Clarifying the input helps the supervised encoder

Table 1. Encoder ablation on 66 human-consensus calls, including 17 positives. These
classifiers differ from the production classifier in Table 3.

Input representation

STOP F1 ROC-AUC

Raw transcript
0.7879
Normalized statements 0.8235

6

0.9304
0.9616
Page 7 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 7 of 21
Text of page 7
Under review as a conference paper at ICLR 2027

Normalization increases both reported metrics without selecting evidence. The larger silverlabel comparison is consistent in direction: precision increases by 9.8, 11.9, 15.6, and 15.8
percentage points at the reported recall settings of 0.60, 0.70, 0.80, and 0.90, respectively
(Appendix B.2). Those gains measure improved teacher agreement; the human pilot supplies
independent but less precise evidence.

5.2

Focusing the input further helps several prompted readers

Table 2. STOP F1 for nine prompted readers on the same 66 calls. Raw is the raw transcript,
Full norm. the complete normalized call, and Selected norm. the subset of normalized statements chosen by the subject-match selector; the final column is the F1 difference between
selected and raw input. Scores are from the 66-call pilot, so per-reader differences are small
and mixed in direction.

Reader

Full norm. Selected norm. ∆(sel, raw)

glm-4.7-flash
0.4545 0.5833
nova-micro
0.4516 0.5405
nemotron-nano-12b 0.3810 0.5185
llama4-scout-17b
0.5833 0.5833
nova-lite
0.6222 0.5957
llama4-maverick-17b 0.7568 0.8125
claude-haiku-4.5
0.8235 0.8571
llama-3.3-70b
0.8000 0.7333
qwen3-32b
0.7000 0.6250

Raw

0.7857
0.7692
0.6364
0.6897
0.7317
0.7895
0.8235
0.7500
0.5614

+0.3312
+0.3176
+0.2554
+0.1064
+0.1095
+0.0327
0.0000
-0.0500
-0.1386

The median paired F1 change is +0.0336 from raw to full input and +0.1064 from full to
selected input. For several lower-scoring readers, both stages help: GLM improves from
0.4545 to 0.5833 to 0.7857, and nova-micro from 0.4516 to 0.5405 to 0.7692. Selection adds
value beyond full normalization in these configurations.

The effect depends on the reader. Haiku performs best on the full normalized call; Qwen
degrades across both transformations. Thus, clearer text and a narrower input are separate
interventions whose usefulness must be checked for the chosen reader. Because selection operates on normalized statements, this comparison does not establish the benefit of rewriting
over an evidence-matched selection of original turns. Appendix B.3 reports the exploratory
cross-reader associations.

5.3

A small producer supports competitive task quality

Replacing teacher normalization with a 0.6B student under an unchanged production encoder yields a serving pipeline with no large-model inference call.

Table 3. Changing the normalizer under the same production classifier. Metrics use the
66-call human-consensus pilot. The classifier differs from the ablation in Table 1.

Input producer

Precision Recall STOP F1 ROC-AUC

Teacher normalizer
0.8750
0.6B student normalizer 0.9286

0.8235 0.8485
0.7647 0.8387

0.9676
0.9556

The student pipeline reaches 0.8387 F1, exceeding the best reported configurations of Llama
3.3 70B (0.8000) and Maverick (0.8125), as well as Haiku on raw or selected input (0.8235).
Haiku on full teacher-normalized input achieves 0.8571, the highest prompted F1 in the
reported sweep. We therefore use that configuration as the quality-relevant cost reference
in Section 6.

Student substitution increases precision and reduces recall. The F1 change is -0.0098, with
a reported bootstrap interval of [-0.0544, +0.0185], which includes both zero and practically
relevant loss. At the reported 0.90 recall setting, precision is 0.5484 with student input
versus 0.8421 with teacher input. This high-recall gap remains material despite the close

7
Page 8 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 8 of 21
Text of page 8
Under review as a conference paper at ICLR 2027

headline F1; the pilot’s coarse recall steps require the operating points to be interpreted
through their achieved recall and thresholds.

Student output contains approximately 11% more statements and 9% more tokens per call,
without establishing the cause of the quality change. Across three seeds per size, within-size
variation exceeds the small 0.6B-versus-1.7B difference. The experiments therefore support
a working small-model substitution, with quality assessed at the intended operating point.

5.4

Zero-shot transfer of a reusable decision reader

A reusable reader could answer new questions from their definitions, reducing the need to
train and maintain a separate classifier for each task. We evaluate the released Laya typeddecisions checkpoint (Convai Innovations, 2026b), an open-weight Jev-compatible reader,
on selected normalized statements without STOP-specific training. Across four compact
task definitions, ROC-AUC ranges from 0.607 to 0.733, compared with 0.9616 for the tasktrained encoder on the full normalized call. The released checkpoint does not recover the
supervised reader’s quality on this policy. Training, input scope, and definition length differ;
Appendix D reports the protocol and all four variants.

6

Quality and cost at scale

At twenty questions per conversation, the reported rates give a projected serving cost of
about $398 per million calls for the student pipeline, including normalization, versus $38,319
for Haiku’s reader alone. This is a 96.2-fold difference, or approximately 99% lower cost.
The projection combines inexpensive normalization with cheap recurring encoder decisions.

Let B be the shared cost of normalization and required metadata, c q the inference cost for
question q, and n the number of distinct questions. The serving cost is

C(n) = B +

n
∑

c q .

q=1

The reported rates per million calls are a student normalization cost B S = $264.85, $6.68
per encoder pass, and $1,915.95 per question for Haiku reading the full normalized call.
These are conditional estimates on the reported offline batch/list-price basis, not invoiceverified spend. The student rate is paired with its measured F1 of 0.8387. Haiku’s 0.8571
uses teacher-normalized input, whose separate construction cost B T remains unresolved:

C student (n) = B S + 6.68n,

C Haiku (n) = B T + 1,915.95n.

Table 4. Observed STOP F1 and projected serving cost per million calls (dollars). Student
costs include normalization; Haiku costs exclude B T . n > 1 projects separately scored
questions at the stated rates, not measured multi-task quality. Costs are rounded.

Serving configuration

STOP
F1

n = 1 n = 20 n = 50

0.6B normalizer + encoder
Haiku on full teacher-normalized input, reader only

0.8387
0.8571

272
1,916

398
38,319

599
95,798

The ratio is 7.1 for one question and 160.0 for fifty questions. Adding Haiku’s teachernormalization cost, B T , increases these ratios. The measured quality comparison is 0.8387
versus 0.8571 STOP F1; the multi-question workloads project separate task passes at the
stated rates.

The comparison concerns serving under separate task passes. It excludes labeling, training,
calibration, and maintenance. Multi-question prompts, provider caching, shared encoder
heads, summaries, and retrieval indexes are alternative reuse strategies that need their own

8
Page 9 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 9 of 21
Text of page 9
Under review as a conference paper at ICLR 2027

quality and cost measurements. Rates must also cover the evaluated input lengths and
any required metadata production. Appendix C gives the detailed projections, lower-priced
baselines, and accounting assumptions. The result supports a conditional cost advantage at
nearby observed F1, not universal cost dominance.

7

Discussion and limitations

Evidence and scope. The human pilot covers one task in one English customer-service
domain and cannot establish narrow performance differences or transfer to new tasks. Consensus filtering can exclude difficult cases. The larger evaluation measures agreement with
the same teacher that creates the representations. Reuse is already exercised at the tagging
stage, where three classification heads share one encoder trunk over the same stored statements. Adding further downstream questions follows the same pattern: a new selector and
head over an unchanged representation. We have not yet measured end-to-end quality for
several downstream questions at once.

What the intervention changes. Rewriting, canonicalization, tagging, and selection are coupled by design. Selection is only as reliable as the tags, and tags are only as reliable as the
units they label. Tagging raw turns would attach one label to several mixed claims, so normalization is what makes tag-based selection possible. The coupling still leaves each stage’s
separate contribution open. An evidence-matched original-turn control, surface cleanup,
summary, and raw-text retrieval baselines are needed to isolate their contributions. Normalization can lose qualifications and selection can discard necessary evidence; source fidelity
and tagging correctness require independent evaluation. Reuse is limited to information the
representation retains.

Can one decision model serve many questions? Task-trained heads currently provide the
stronger STOP results, while reusable decision readers could reduce the labeling, calibration, and maintenance work that recurs with each new question. The stored representation
supports either route. A reusable reader must recover the decision rule from its definition as
well as interpret the evidence. Further comparisons should measure the benefit of normalization and selection for reusable readers and weigh any quality loss against the reduction
in per-question setup.

8

Conclusion

Conversation analytics repeats the same interpretive work every time a new question is asked
of an old transcript. We propose doing that work once. Statement normalization turns each
conversation into short, attributed, tagged statements, stored for later questions. For each
question, a downstream reader then consumes either all statements or a slice chosen by a
Boolean predicate.

On the offer-suppression task, normalization improves the supervised encoder without selection, while several weaker prompted readers benefit from both clearer statements and
a narrower evidence set. A 0.6B normalizer supplies a fixed production encoder at 0.8387
F1, close to the best reported prompted configuration’s 0.8571. At the reported rates, this
pipeline has approximately 100-fold lower projected serving cost at twenty questions per
conversation. Laya’s zero-shot results identify a remaining challenge: extending the strong
task-trained results to a reader that accepts new decision definitions without a separate
trained head.

The human evidence covers one business policy and a small expert-labeled cohort; multiquestion savings are projected. Within this setting, the results support a practical division
of work: make meaning explicit once, select evidence when it helps, and use small models for
recurring decisions. Evaluating more questions over the stored representation will establish
how far that reuse extends.

9
Page 10 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 10 of 21
Text of page 10
Under review as a conference paper at ICLR 2027

9

Ethics statement

The study analyzes enterprise customer-service calls, and the evaluated label concerns suppression of further product offers. Errors can change which customers continue to receive
offers. Aggregate agreement alone does not establish that two systems affect the same individuals, so operational use requires calibration and review at the intended decision threshold.

Access to the underlying transcripts is restricted by the deployment’s data agreement. Brand
canonicalization and speaker attribution are representation operations; they should not be
interpreted as guarantees of anonymization. The example is an attributed agent–customer
exchange. Its statements record what the speakers said, including the agent’s offer and the
customer’s plans, without independently verifying those claims.

10

Reproducibility statement

The main tables distinguish encoder ablation, prompted-reader inputs, and fixed-classifier
normalizer substitution. Appendices B–D report supplementary analyses, cost assumptions,
and the exploratory Laya protocol. The private transcript corpus is not publicly released,
limiting exact replication. Reimplementation on another corpus would test transfer rather
than reproduce these particular scores.

11

AI-use statement

AI assistance was used to revise the paper’s narrative, review methodological choices and
interpretations, check consistency among reported quantities, and propose further analyses.
Language models were also used within the experimental pipeline to construct representations and supervision and to make prompted predictions. The authors are responsible for
verifying the final text, results, citations, and disclosures against the underlying research
records.

A

Appendix A. Extended related work

Our pipeline makes three commitments, and each one meets a separate line of prior work.
The first concerns the unit into which a conversation is rewritten: short statements that
carry one claim, a speaker, a source reference, and a qualification. The second concerns how
those units reach a reader: either all at once or through a Boolean predicate over controlled
tags. The third concerns when interpretation is paid for and by which model: once per
conversation, before any question is known, by a small distilled producer. Much of each
commitment has precedent. This appendix follows those precedents in turn, noting what
each line of work established, the setting it was built for, and the part of the problem it
leaves open. The final subsection states which combination we believe is new.

A.1

A.1 From sentences to self-contained units

The idea that text becomes easier to use once it is broken into small, self-contained pieces
has a long history. Split and Rephrase framed the task directly: rewrite a complex sentence
as a meaning-preserving sequence of shorter sentences, supported by a benchmark of over
a million complex-to-simple examples (Narayan et al., 2017). Decontextualization added
a second requirement. A sentence taken from a Wikipedia article should be rewritten so
that it can be interpreted without its surrounding context, while keeping its meaning (Choi
et al., 2021). Splitting alone can lose the relations between the pieces, and discourse-aware
simplification addresses this by producing a hierarchy of core and context sentences linked
by rhetorical relations (Niklaus et al., 2022). Open information extraction reached a similar
concern from the opposite direction. MinIE shortens relational tuples but moves polarity,
modality, attribution, and quantities into explicit annotations, so that minimizing a fact
does not silently change what it asserts (Gashteovski et al., 2017).

10
Page 11 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 11 of 21
Text of page 11
Under review as a conference paper at ICLR 2027

Recent work combines these requirements in the atomic proposition. Dense X Retrieval
defines a proposition as a distinct, minimal, and self-contained piece of meaning, and indexes
English Wikipedia as roughly 257 million of them (Chen et al., 2024). Its Propositionizer
is trained by distillation: GPT-4 decomposes 42k passages, and a Flan-T5-large model is
fine-tuned on the result. FActScore decomposes generated text into atomic facts in order to
measure factual precision (Min et al., 2023).

This line of work supplies the design principles our statements inherit, and Dense X shows
that a large model’s decomposition can be taught to a small one. The text it was developed
on, however, is written, monologic, and largely declarative. A conversation adds problems
these methods were not built for. Every claim has a speaker, and the same words mean
different things from an agent and from a customer. Positions change within a call, so
that a deferral in one turn and a conditional acceptance in the next are both part of the
customer’s stance. Many claims are commitments or intentions rather than facts, and their
conditions, such as “after I move,” carry the business meaning. Our statements therefore
add speaker attribution, links back to source turns, and an explicit qualification tag to the
proposition principles. They are judged by whether downstream decisions improve, rather
than by retrieval granularity or factual precision.

A.2

A.2 Rewriting dialogue for the next component

Dialogue research arrived at rewriting through a different need. In a preliminary study of
2,000 Chinese multi-turn conversations, Su et al. (2019) found coreference or omission in
more than 70% of utterances. They collected 20k annotated dialogues, trained a pointerbased Transformer to rewrite each utterance so that it recovers referenced and omitted
content, and integrated the rewriter into two online chatbots, where it improved intention
detection and user engagement. Rastogi et al. (2019) treated reference resolution in multidomain spoken dialogue as context-aware query reformulation, so that downstream languageunderstanding components receive a self-contained request.

These systems show that a rewritten conversation is a better interface for downstream
models than the raw one, which is also our premise. Their rewriting is prospective and
local, however. It runs online, one turn at a time, and serves a single question: what does
the user want now, so that the system can act next. Our normalization is retrospective
and global. It runs after the call has ended, over the whole conversation, and produces an
inventory of every attributed claim for questions that no one has asked yet.

Simplification has also been tested as preprocessing for specific downstream tasks. Simple-
SQuAD rewrites SQuAD contexts with a style-transfer simplifier, locates each answer span
again in the simplified text, and reports gains of up to 2.04 exact-match points (Dadu et al.,
2021). DisSim-FinBERT applies discourse simplification to Federal Open Market Committee minutes, improving aspect selection and sentiment prediction (Kim et al., 2025). Both
results confirm that simpler units can help a reader, but in each case the transformation is
built and validated around one known task. Simple-SQuAD, in particular, repairs answer
positions after simplification, which is possible only because the answers are known in advance. A representation for analytics has to preserve evidence for answers that do not yet
exist.

A.3

A.3 Structure that makes text addressable

Labeling what an utterance does is an established part of dialogue analysis. The ISO
24617-2 standard annotates functional segments with communicative functions across several dimensions, together with qualifiers for properties such as certainty, conditionality, and
sentiment (Bunt et al., 2017). Our speech-act and qualification tags echo this design. The
standard, however, is a general scheme for describing dialogue. Our tags form a small,
application-specific vocabulary whose subject family names business topics, and their purpose is operational: they exist so that a rule can select evidence.

The closest prior proposal is Telegraph English (Arbuzov et al., 2026). It rewrites text,
without reference to any question, into atomic fact lines that use about forty logical and

11
Page 12 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 12 of 21
Text of page 12
Under review as a conference paper at ICLR 2027

relational symbols. It tags lines for temporal state, modality, semantic role, and scope, and
argues that each line becomes an independently addressable unit that a query can retrieve
together with its scope. Its evaluation measures the rewrite as static compression, over 4,081
LongBench-v2 questions and five OpenAI models, at compression ratios matched against
LLMLingua-2. At about half the original tokens, it preserves 99.1% of key-fact accuracy
with GPT-4.1, and its advantage grows for smaller models, reaching 11 points on fine-detail
questions. Its authors state that dynamic context management is not yet benchmarked, and
they describe compress-once reuse as an architectural argument rather than an empirical
result.

Our study begins where that evaluation stops. We measure selection as its own intervention,
comparing full and selected representations within the same readers. The representation
targets conversational register rather than general documents. We add a supervised encoder
that benefits without selection, and we distill the producer. We also do not treat length as
the objective. Normalization is not designed to shorten a call, and its intended effect is on
what the reader must infer rather than on how many tokens it reads.

Structured memory for conversational agents is a second neighbor, and it also operates on
dialogue. A-MEM turns each interaction into a note with a contextual description, keywords,
and tags. It links related notes, updates older notes as new ones arrive, and retrieves the topk notes by embedding similarity. It is evaluated on LoCoMo, whose conversations average
about 9,000 tokens across as many as 35 sessions (Xu et al., 2025). SimpleMem compresses
interactions into compact units at write time, synthesizes related units into more abstract
ones, indexes them under several views, and plans retrieval from the inferred intent of each
query. It reports a 26.4% average F1 gain on LoCoMo with up to thirty times fewer inference
tokens (Liu et al., 2026).

These systems answer a different question: what should one agent remember about its own
continuing relationship with a user. Their memory is mutable and consolidated over time,
and access is based on similarity or planned by a language model for each query. Our setting
is an analyst’s view over millions of completed conversations between other people. Each
call’s representation is written once and not revised. Access is a deterministic predicate over
a fixed vocabulary, so the same slice can be reproduced across a corpus and audited against
the source turns, and no model call sits in the selection path.

A.4

A.4 Reducing what the reader must read

Context compression treats the reader’s input length as the cost to reduce, and its main
division is whether the compressor sees the question. RECOMP trains extractive and abstractive compressors on the reader’s end-task performance. Its compressors can return an
empty string when retrieved documents do not help, and they compress to about 6% of the
original length with little loss (Xu et al., 2024). AttentionRAG reformulates each query
so that its focus falls on a single token, then keeps the context that this token attends to,
reaching up to 6.3 times compression (Fang et al., 2025). Both methods judge relevance
with the current question in hand, so the cost of compression recurs with every question.

Question-independent compression avoids that recurrence. LLMLingua-2 distills GPT-4
compression decisions into a token classifier built on a small encoder, and deletes tokens
without seeing the task (Pan et al., 2024). One of its results anticipates a pattern in ours.
With Mistral-7B as the reader, the compressed MeetingBank prompt scored higher than
the original, 76.22 against 66.95 exact match, which the authors attribute to the smaller
model’s weaker handling of long contexts. ReadAgent splits a long document into pages
and compresses each page into a gist before the question is shown. When a question arrives,
it looks up the original pages it needs, extending effective context by 3.5 to 20 times on
QuALITY, NarrativeQA, and QMSum (Lee et al., 2024).

These methods establish that weaker readers can benefit from shorter inputs, and that
question-independent preparation combines well with question-time lookup. Our selection
stage shares that structure but differs in what it operates on and how it decides. Deletion
leaves no structure to query, and gists fix in advance which details survive. Our selection
operates on a representation that keeps every attributed claim and labels it, so relevance

12
Page 13 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 13 of 21
Text of page 13
Under review as a conference paper at ICLR 2027

becomes a rule rather than a model judgment. Normalization and selection also appear as
separable effects in our results. Compression work does not separate them, because it makes
the input clearer and shorter in the same operation.

A.5

A.5 Paying for interpretation once

Doing interpretive work before questions arrive predates neural models. Fleischman et al.
(2003) mined about two million concept–instance relations offline from 15GB of newspaper
text, filtered them with a learned classifier, and answered “Who is” questions from the
resulting repository. Compared with a web-based question-answering system, the repository answered 25% more questions correctly and did so three orders of magnitude faster.
EVAPORATE brings the same logic to language models. It builds structured tables from
semi-structured documents, often by having the model write extraction functions instead of
reading every document, and reduces the tokens the model must process by a factor of 110
on average across 16 settings of 10k documents each (Arora et al., 2023).

Index-building retrieval systems apply the idea at corpus scale. GraphRAG extracts an
entity graph and precomputes summaries of graph communities, independently of any query,
in order to answer global questions over corpora of about a million tokens (Edge et al.,
2024). RAPTOR recursively clusters and summarizes text into a tree, and raises the best
reported accuracy on QuALITY by 20 absolute points when paired with GPT-4 (Sarthi et al.,
2024). Within conversation analytics, the most natural question-independent representation
is a summary, and TWEETSUMM provides about 6,500 human-written extractive and
abstractive summaries of customer-care dialogues (Feigenblat et al., 2021).

All of these approaches pay once, and each fixes in advance what the prepared representation keeps. Fleischman’s repository keeps one relation type. EVAPORATE’s tables keep
the attributes that were extracted. Summaries and community reports keep what the summarizer judged salient, and GraphRAG and RAPTOR still call a large model at query time.
Our representation fixes the unit and the tag vocabulary, but not the set of facts. Every
claim remains in the stored call as natural language, so a question nobody anticipated can
still be answered from it, and the query path contains only a rule and a small encoder. A
summary remains the most important untested alternative in our setting, and Section 7 lists
it among the required baselines.

A.6

A.6 Small models for the recurring cost

The last commitment is economic: the recurring stages should run on small models.
Sequence-level knowledge distillation trains a student on the teacher’s generated outputs
rather than on reference outputs (Kim & Rush, 2016), and our normalizer student is trained
this way. Distilling step-by-step adds teacher rationales as extra supervision, so that a 770M
T5 model outperforms few-shot 540B PaLM on one benchmark using 80% of the training
data (Hsieh et al., 2023). FrugalGPT reduces cost at query time instead, by sending each
query through a learned cascade of models and stopping once an answer is reliable, which
matches GPT-4 with up to 98% cost reduction (Chen et al., 2023). Ask Me Anything shows
that reformatting inputs can close much of the gap between model sizes. The model rewrites
each input into several open-ended question-answer prompts, the answers are aggregated,
and GPT-J-6B matches or exceeds few-shot GPT-3-175B on most benchmarks tested (Arora
et al., 2022).

In each case the saving attaches to a single task or a single query. A distilled task model
serves the task it was trained for, a cascade still makes at least one model call per question,
and Ask Me Anything asks the reader to reformat its own input every time. Our pipeline
distills the one component that all questions share. A single distillation serves every later
question, and the reformatting that Ask Me Anything performs inside the reader happens
once, upstream, in a different model. A weak reader therefore receives structured input
without having to produce it.

13
Page 14 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 14 of 21
Text of page 14
Under review as a conference paper at ICLR 2027

A.7

A.7 What is and is not new

Most individual ingredients have precedent: self-contained units, dialogue rewriting,
dialogue-act tags, question-independent preparation, distilled producers, and small readers that benefit from denser input. We do not claim any of them. The contribution lies
in their combination for conversation analytics and in its measurement. A conversation is
normalized once into attributed, qualified, source-linked statements. Controlled tags then
turn evidence selection into a rule rather than a model call. We measure normalization
and selection separately, across a supervised encoder and nine prompted readers, against
human labels. Finally, a distilled 0.6B producer is substituted under a fixed classifier, which
removes large-model calls from the serving path.

B

Appendix B. Representation and supplementary analyses

B.1

B.1 Statement structure and programmatic slices

Example 1 separates three subjects within customer turn 92: deferring the phone line, moving home, and requesting TV setup. The canonical brand tokens expose explicit mentions,
while subject tags also retrieve responses that do not repeat a brand. The final statement
preserves both future intent and its condition, “After I move.”

Source references and statement order remain useful. References to “the phone line” depend
on the earlier offer, and 88.2 retains two related claims about the brands and customer
relationship. Self-contained statements are a design objective; source fidelity and preserved
qualifications take precedence over forcing every output into an isolated sentence.

The following predicates select statements from Example 1 in the Introduction. contains
denotes a literal match against statement text; act, obj, and mod refer to the appended
tags. Results retain source order.

Table B1. Exact selections from the running example.

Evidence slice

Explicit Boost Mobile contains(”[BOOST_MOBILE]”)
88.1, 88.2
mentions
Explicit mentions of both contains(”[BOOST_MOBILE]”) AND con- 88.2
brands
tains(”[DISH]”)
Wireless discussion
obj = WIRELESS
88.1, 88.2, 89.1, 90.1, 91.1,
92.1, 93.1
Agent’s wireless offers
speaker = A AND act = offer AND obj = 88.1, 89.1
WIRELESS
Customer’s wireless de- speaker = C AND act = refuse AND obj = 92.1
ferral
WIRELESS
Customer’s
intended speaker = C AND act = accept AND obj 93.1
wireless acceptance
= WIRELESS AND mod = intent
Customer’s
wireless speaker = C AND obj = WIRELESS AND 92.1, 93.1
stance (Figure 1)
(act = refuse OR act = accept)
Move and TV setup
obj = MOVE OR obj = TECH_VISIT
92.2, 92.3

Predicate

Selected statement IDs

These slices illustrate access to the same stored representation. They are not separate endto-end task evaluations. A task concerning the customer’s stance may need both 92.1 and
93.1; a narrow predicate is useful only when it retains the evidence required by the decision.

B.2

B.2 Encoder precision at the reported recall settings

Table B2. Normalized-minus-raw precision on 16,050 held-out calls. Both input representations fit the context cap; reference labels are teacher-generated.

14
Page 15 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 15 of 21
Text of page 15
Under review as a conference paper at ICLR 2027

Recall setting Precision difference

0.60
0.70
0.80
0.90

+0.098
+0.119
+0.156
+0.158

The normalized arm leads at all four reported settings. Because the teacher supplies both
the normalized text and the reference labels, this comparison measures improved teacher
agreement. It does not replace the independent human evaluation in Table 1 or isolate
rewriting from deterministic entity canonicalization.

B.3

B.3 Exploratory reader association

Using the rounded values in Table 2, the Spearman correlation between raw F1 and selectedinput F1 is 0.4167. The correlation between raw F1 and the change from raw to selected
input is -0.8000.

The second statistic is mathematically coupled: the raw score is also subtracted when computing the gain. For Pearson covariance, the corresponding identity is

Cov(X, Y − X) = Cov(X, Y ) − Var(X).

Although this identity does not calculate Spearman correlation, it illustrates why variation
and measurement error in the baseline can contribute to an inverse association. The result is
therefore descriptive rather than evidence identifying a capability-substitution mechanism.

An exploratory split into four lower- and four higher-raw-F1 readers, excluding the median
reader, yields group medians of 0.45305 and 0.77840 on raw input, and 0.72945 and 0.76975
on selected input. The gap between those group medians decreases from 0.32535 to 0.04030.
This is distinct from the range across all nine readers, which decreases from 0.4425 to 0.2621.

B.4

B.4 Smaller training-family results

A separate experiment family based on 7,305 calls produced normalized-input and selectedinput classifiers whose performance on the human pilot was near chance. The paired interval
for that representation comparison included zero. A student-versus-teacher comparison in
the same experiment family favored student input, with a reported interval excluding zero,
but involved only a few calls and did not resolve the underlying poor absolute performance.

These outcomes provide no basis for attributing the problem solely to the evaluation set.
They may reflect training, representation, calibration, sampling, or other differences in that
experiment family. We therefore keep them separate from the production-classifier result in
Table 3 and do not use them to establish student superiority.

B.5

B.5 A synthetic call, worked end to end

Example 2. The exchange below is synthetic, not drawn from the corpus. It shows one
bounded window carried through the whole construction: the raw recognizer transcript
with its global turn indices, then the normalized, tagged statements. Every transformation
named in Section 3.1 appears — the greeting at turn 1 is dropped, disfluencies and false
starts are removed, references are resolved, brands are canonicalized, the redaction token is
copied verbatim, and each speaker’s stance is kept.

• [1] A: Hi, thanks for holding, this is Nicole.

• [2] C: uh yeah, hi, this is Denise Holloway.

• [5] C: I wanted to add the, uh, the sports package, for the college football this
weekend.

• [10] A: sure, so it’s uh $13.00 a month, but I can do 50% off, so that’s $6.50.

15
Page 16 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 16 of 21
Text of page 16
Under review as a conference paper at ICLR 2027

• [11] C: OK, yeah, that, that sounds great, let’s do it.

• [15] A: and will you still be using the card ending in [CREDIT_DEBIT_NUMBER]
for autopay?

• [18] A: and, oh, are you familiar with our cell phone service, Boost Mobile?

• [19] C: no, I’m good for now.

Dropping the greeting at turn 1, normalization yields:

2.1 C My name is [DENISE_HOLLOWAY]. inform IDENTITY actual
5.1 C I want to add the sports package for the college football games this
weekend. request SPORTS_PACKAGE actual

10.1 A The [MULTI_SPORT_PACK] is $13.00 a month, but I can add it at
50% off, which is $6.50. offer SPORTS_PACKAGE actual
11.1 C Please add the
[MULTI_SPORT_PACK]. accept SPORTS_PACKAGE actual

15.1 A Will you still be using the card ending in [CREDIT_DEBIT_NUMBER]
for autopay? ask PAYMENT_METHOD actual
18.1 A Are you familiar with our cell phone service,
[BOOST_MOBILE]? offer WIRELESS actual
19.1 C I am not interested in [BOOST_MOBILE] right
now. refuse WIRELESS actual

The generic ”sports package” at turn 5 names no product and is left untagged for entity,
while the agent’s specific offer at turn 10 canonicalizes to [MULTI_SPORT_PACK]; ”let’s
do it” at turn 11 is read as an acceptance only because turn 10 established the offer, and the
customer’s closing ”no, I’m good for now” keeps its force as a refusal rather than softening
into a deferral.

C

Appendix C. Cost projections and accounting assumptions

C.1

C.1 Component rates and full projection

The main comparison uses the strongest prompted configuration, Haiku on full teachernormalized input. B T is its teacher-normalization cost per million calls; it is left explicit
because the supplied rates do not resolve that stage. The student pipeline’s $264.85 construction rate is paired only with the student-input classifier’s F1 of 0.8387, not the teacher-input
classifier’s 0.8485. Haiku’s quality on student input has not been measured.

Table C1. Expanded quality and cost comparison. Costs are dollars per million calls,
rounded; n > 1 projects separate task passes. Add B T to every Haiku amount.

STOP n = n = n = n =
F1
1
5
20
50

Serving configuration

0.6B normalizer + encoder, including normalization
0.8387 272 298 398
599
Haiku 4.5 on full teacher-normalized input, reader cost only 0.8571 1,916 9,580 38,319 95,798

Before rounding, student totals are $271.53, $298.25, $398.45, and $598.85. Haiku readeronly totals are $1,915.95, $9,579.75, $38,319.00, and $95,797.50. Ratios in Section 6 use these
amounts. The reported marginal encoder pass costs approximately 1/287 of the Haiku pass.

C.2

C.2 Lower-priced prompted baselines

The original Nova cost projections remain useful context for the broader quality–cost tradeoff. Their lower F1 scores make them less informative reference points for the main comparison in Section 6.

16
Page 17 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 17 of 21
Text of page 17
Under review as a conference paper at ICLR 2027

Table C2. Conditional costs for lower-priced raw-input readers. Costs are dollars per million conversations, rounded to the nearest dollar. F1 is measured only on STOP; n > 1
projects separately scored tasks at the stated rates. The student row includes its assumed
normalization cost.

Configuration and cost formula

0.6B normalizer + encoder: 264.85 + 6.68n 0.8387
nova-micro on raw input: 82.41n
0.4516
nova-lite on raw input: 142.21n
0.6222

STOP F1 n = 1 n = 5 n = 20 n = 50

272
82
142

298
412
711

398
1,648
2,844

599
4,121
7,111

Under these rates, the student pipeline becomes cheaper than raw-input nova-micro at the
fourth task. This crossover concerns alternatives with substantially different observed task
quality. The main comparison instead uses Haiku’s strongest evaluated input configuration.

C.3

C.3 What the projection includes

∑
For a fixed corpus, the general model B + q c q allows different selectors, input lengths,
and readers to have different marginal costs. The tabulated scenarios hold a per-task rate
fixed to illustrate amortization. They estimate serving and exclude development costs such
as labels, training, calibration, and schema design.

The normalization and encoder rates must correspond to the evaluated student configuration, including its longer output and any required metadata-generation stage. The reported
batch/list-price estimates require reconciliation with hardware or provider, throughput, numerical precision, batch size and concurrency, utilization, token distributions, price date,
and discounts. The stated totals are conditional until those inputs are resolved against the
run and pricing records.

Alternative serving strategies can change the comparison. Multiple questions can share a
prompt; caching can reduce repeated input processing; raw-transcript encoders can share
computation across heads. Summaries and retrieval indexes provide other reusable representations. Their costs are informative only alongside quality on the same tasks.

Caching one pooled whole-call encoder vector is also a different computation from reencoding a question-specific selected span. The tag-head experiments do not establish equal
quality for that cached variant. Multi-task evaluation should therefore measure representation reuse, reader reuse, and per-question setup costs separately.

D

Appendix D. Zero-shot Laya evaluation

D.1

D.1 Protocol

We evaluate the released Laya typed-decisions checkpoint on selected normalized statements
from the same 66 human-consensus calls, including 17 STOP positives. Zero-shot here
means no additional training, fine-tuning, or labeled STOP demonstrations for this task;
the released checkpoint already has its own prior training. We do not evaluate Laya on the
larger silver-label corpus.

The runner uses max_len = 1024 and head_max_len = 256, with each answer option
capped at 48 tokens. The original wording uses approximately 165 head tokens: 77 for
instructions and 88 for option criteria. It is a compressed task specification, not the complete
approximately 1,000-token rubric used by the generative readers. Follow-up variants enrich,
shorten, or rephrase that specification within the budget.

The initial run reports two truncated inputs out of 66, both non-STOP, and no empty
selections. All 17 positive calls remain in the denominator. This rules out truncation of
positive calls in that run; it does not establish that the selected text retained all necessary
evidence or that truncation had no effect on negative-call scores.

Each wording is evaluated once with fixed released weights. This is a wording-sensitivity
analysis, not a repeated-run or training-seed study. The follow-up wordings were evaluated

17
Page 18 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 18 of 21
Text of page 18
Under review as a conference paper at ICLR 2027

after the initial result was known, so the analysis is exploratory. We report every variant
rather than presenting the highest-scoring wording as a new independent confirmation.

D.2

D.2 Results and interpretation

Table D1. Laya wording sensitivity on the 66-call human cohort. Laya precision, recall, and
F1 use threshold 0.5. ROC-AUC uses the returned STOP scores. The supervised reference
is the full-normalized-input ablation classifier in Table 1, not either production-classifier row
in Table 3. The comparison shares labels and calls, but differs in training and input scope.
Values retain the precision supplied in the experiment report.

Reader / wording

STOP precision

Recall

F1

ROC-
AUC

Laya: q1_original
Laya: q2_detailed, enriched definition
Laya: q3_terse, two sentences
Laya: q4_reworded
Task-trained encoder, full normalized input

0.556
0.333
0.429
0.323
0.8235

0.294
0.118
0.529
0.588
0.8235

0.385
0.174
0.474
0.417
0.8235

0.643
0.607
0.733
0.678
0.9616

All four Laya variants underperform the supervised reference in both reported metrics. The
detailed definition performs worst in this probe; that observation does not establish that
richer task definitions are generally ineffective. F1 changes substantially at the fixed threshold, and ROC-AUC also changes by 0.126 across wordings. The differences therefore cannot
be attributed entirely to threshold crossings.

ROC-AUC measures ranking without selecting an operating threshold. A threshold change
or monotone score recalibration would not remove the observed ranking gap. We do not
infer that the four wordings are statistically equivalent: no paired uncertainty estimate
is supplied for their differences. The experiment also does not establish that fine-tuning
is necessary or sufficient to close the gap. Model choice, task-definition fidelity, evidence
selection, and adaptation remain separate factors for future study.

D.3

D.3 Implications for reusable readers

A task-trained classifier learns a decision boundary from labels, whereas a zero-shot reader
must obtain it from the question and definition alongside the evidence. At hundreds of
questions, reusable schema-driven readers could reduce the recurring work of fitting, validating, calibrating, and maintaining separate heads. The present experiment leaves that
opportunity open while showing that this checkpoint and these definitions do not recover
the STOP classifier’s quality.

Further comparison should hold task definitions and evidence inputs fixed, testing raw,
full normalized, and selected text where context budgets permit. Zero-shot use, limited
adaptation, and task-trained heads should be compared at explicit quality requirements,
with per-question setup costs reported alongside inference. Wordings should be developed
away from the final evaluation set.

References

Diogo Almeida. Introducing System One Models & Jev. TypeSafe AI Blog, September
2026. URL https://typesafe.ai/blog/introducing-system-one-models-and-jev. Published
September 15, 2026. Accessed September 25, 2026.

Mikhail L. Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev, and Alexey A. Shvets. Telegraph English: Semantic prompt compression via structured symbolic rewriting. arXiv
preprint arXiv:2605.04426, 2026. doi: 10.48550/arXiv.2605.04426.

Simran Arora, Avanika Narayan, Mayee F. Chen, Laurel Orr, Neel Guha, Kush Bhatia, Ines
Chami, and Christopher Ré. Ask me anything: A simple strategy for prompting language
models. In International Conference on Learning Representations (ICLR), 2022.

18
Page 19 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 19 of 21
Text of page 19
Under review as a conference paper at ICLR 2027

Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel
Trummer, and Christopher Ré. Language models enable simple systems for generating
structured views of heterogeneous data lakes. Proceedings of the VLDB Endowment, 17
(2):92–105, 2023. doi: 10.14778/3626292.3626294.

Sisong Bei, Mikhail L. Arbuzov, Ziwei Dong, Dmitri Kalaev, and Alexey Shvets. Context
compression is not one thing: Readable symbolic re-expression vs. coherent summary at
matched budget. arXiv preprint arXiv:2606.14875, 2026. doi: 10.48550/arXiv.2606.14875.

Harry Bunt, Volha Petukhova, David Traum, and Jan Alexandersson. Dialogue act annotation with the ISO 24617-2 standard. In Deborah A. Dahl (ed.), Multimodal Interaction with W3C Standards, pp. 109–135. Springer, Cham, 2017. doi: 10.1007/
978-3-319-42816-1_6.

Lingjiao Chen, Matei Zaharia, and James Zou. FrugalGPT: How to use large language
models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176,
2023.

Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, and Dong Yu. Dense X retrieval: What retrieval granularity should we
use? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 15159–15177. Association for Computational Linguistics, 2024. doi:
10.18653/v1/2024.emnlp-main.845.

Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and
Michael Collins. Decontextualization: Making sentences stand-alone. Transactions of the
Association for Computational Linguistics, 9:447–461, 2021. doi: 10.1162/tacl_a_00377.

Convai Innovations. Laya. Hugging Face model card and model release, 2026a. URL
https://huggingface.co/convaiinnovations/laya. Accessed 22 September 2026.

Convai Innovations. Laya Typed-Decisions. Hugging Face model card, 2026b. URL https:
//huggingface.co/convaiinnovations/laya-typed-decisions. Accessed September 25, 2026.

Tanvi Dadu, Kartikey Pant, Seema Nagar, Ferdous A. Barbhuiya, and Kuntal Dey. Text
simplification for comprehension-based question-answering. In Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021), pp. 1–10. Association for
Computational Linguistics, 2021. doi: 10.18653/v1/2021.wnut-1.1.

Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven
Truitt, and Jonathan Larson. From local to global: A graph RAG approach to queryfocused summarization. 2024.

Yixiong Fang, Tianran Sun, Yuling Shi, and Xiaodong Gu. AttentionRAG: Attention-guided
context pruning in retrieval-augmented generation. arXiv preprint arXiv:2503.10720, 2025.
doi: 10.48550/arXiv.2503.10720.

Guy Feigenblat, R. Chulaka Gunasekara, Benjamin Sznajder, Sachindra Joshi, David Konopnicki, and Ranit Aharonov. TWEETSUMM - a dialog summarization dataset for customer
service. In Findings of the Association for Computational Linguistics: EMNLP 2021, pp.
245–260, 2021.

Michael Fleischman, Eduard H. Hovy, and Abdessamad Echihabi. Offline strategies for
online question answering: Answering questions before they are asked. In Proceedings
of the 41st Annual Meeting of the Association for Computational Linguistics, pp. 1–7.
Association for Computational Linguistics, 2003. doi: 10.3115/1075096.1075097.

Kiril Gashteovski, Rainer Gemulla, and Luciano Del Corro. MinIE: Minimizing facts in open
information extraction. In Proceedings of the 2017 Conference on Empirical Methods in
Natural Language Processing, pp. 2630–2640. Association for Computational Linguistics,
2017. doi: 10.18653/v1/D17-1278.

19
Page 20 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 20 of 21
Text of page 20
Under review as a conference paper at ICLR 2027

Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling step-by-step!
outperforming larger language models with less training data and smaller model sizes.
In Findings of the Association for Computational Linguistics: ACL 2023, pp. 8003–8017,
2023.

Wonseong Kim, Christina Niklaus, Choong Lyol Lee, and Siegfried Handschuh. DisSim-
FinBERT: Text simplification for core message extraction in complex financial texts. arXiv
preprint arXiv:2501.04959, 2025. doi: 10.48550/arXiv.2501.04959.

Yoon Kim and Alexander M. Rush. Sequence-level knowledge distillation. In Proceedings
of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1317–
1327. Association for Computational Linguistics, 2016. doi: 10.18653/v1/D16-1139.

Kuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John Canny, and Ian Fischer. A humaninspired reading agent with gist memory of very long contexts. In Proceedings of the 41st
International Conference on Machine Learning, volume 235 of Proceedings of Machine
Learning Research, pp. 26396–26415. PMLR, 2024.

Jiaqi Liu, Yaofeng Su, Peng Xia, Siwei Han, Zeyu Zheng, Cihang Xie, Mingyu Ding, and
Huaxiu Yao. SimpleMem: Efficient lifelong memory for LLM agents. arXiv preprint
arXiv:2601.02553, 2026. doi: 10.48550/arXiv.2601.02553.

Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit
Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 12076–12100, 2023.

Shashi Narayan, Claire Gardent, Shay B. Cohen, and Anastasia Shimorina. Split and
rephrase. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2017.

Christina Niklaus, André Freitas, and Siegfried Handschuh. Shallow discourse parsing for
open information extraction and text simplification. In Proceedings of the 3rd Workshop on Computational Approaches to Discourse, pp. 64–76. International Conference on
Computational Linguistics, 2022.

Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei
Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei
Zhang. LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt
compression. In Findings of the Association for Computational Linguistics: ACL 2024,
pp. 963–981. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.
findings-acl.57.

Pushpendre Rastogi, Arpit Gupta, Tongfei Chen, and Lambert Mathias. Scaling multidomain dialogue state tracking via query reformulation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics:
Human Language Technologies, Volume 2 (Industry Papers), pp. 97–105. Association for
Computational Linguistics, 2019. doi: 10.18653/v1/N19-2013.

Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D.
Manning. RAPTOR: Recursive abstractive processing for tree-organized retrieval. In The
Twelfth International Conference on Learning Representations, 2024.

Hui Su, Xiaoyu Shen, Rongzhi Zhang, Fei Sun, Pengwei Hu, Cheng Niu, and Jie Zhou. Improving multi-turn dialogue modelling with utterance rewriter. In Proceedings of the 57th
Annual Meeting of the Association for Computational Linguistics, pp. 22–31. Association
for Computational Linguistics, 2019. doi: 10.18653/v1/P19-1003.

Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said
Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper,

20
Page 21 of Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Page 21 of 21
Text of page 21
Under review as a conference paper at ICLR 2027

Griffin Adams, Jeremy Howard, and Iacopo Poli. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational
Linguistics (Volume 1: Long Papers), pp. 2526–2547. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.acl-long.127.

Fangyuan Xu, Weijia Shi, and Eunsol Choi. RECOMP: Improving retrieval-augmented
LMs with context compression and selective augmentation. In The Twelfth International
Conference on Learning Representations, 2024.

Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-Mem:
Agentic memory for LLM agents. In Advances in Neural Information Processing Systems,
volume 38, pp. 17577–17604. Curran Associates, Inc., 2025. doi: 10.52202/085713-0593.

21