{
  "slug": "answer-accuracy-and-trace-verifiability",
  "title": "Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning",
  "short": "AATV",
  "line": "reasoning",
  "line_name": "Verifiable reasoning",
  "part": null,
  "status": "ICLR 2027 submission",
  "date": "2026-08-25",
  "authors": [
    "Sisong Bei",
    "Mikhail L Arbuzov",
    "Ziwei Dong",
    "Yan Han",
    "Dmitri Kalaev",
    "Yanxin Zhang",
    "Alexey Shvets"
  ],
  "abstract": "Reasoning traces expose intermediate statements, but accepting a trace and obtaining a correct answer are distinct evaluation outcomes. We built a deterministic checker for a small model’s math reasoning traces and found that checkability and correctness come apart: a trace the checker accepts is correct only about one time in three. The model was trained by outcome-only reinforcement learning to write a compact trace language whose equations and checks a program evaluates. We measured answer accuracy and checker acceptance over the same outputs, using a checkpoint panel and seed-paired runs against free-form reasoning at matched token budgets. At matched training-example exposure, the trace interface raises acceptance by about 34 points and lowers accuracy by about 16 points relative to free-form reasoning. The directions of these differences are already present at the first measured checkpoint. Across the model’s own checkpoints, the accuracy–checkability trade-off we pre-registered did not confirm. Checkability and correctness are different quantities and should be measured separately.",
  "tldr": "On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.",
  "pages": 26,
  "html": "https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability/",
  "md": "https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability.md",
  "reader": "https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability/read/",
  "pdf": "https://telegrapher.ai/papers/answer-accuracy-and-trace-verifiability/answer-accuracy-and-trace-verifiability.pdf",
  "arxiv": null,
  "openreview": "https://openreview.net/forum?id=xZMZw6Twv9",
  "post": "https://telegrapher.ai/blog/answer-accuracy-and-trace-verifiability/",
  "bibtex": "@misc{bei2026measuring,\n  title         = {Measuring Answer Accuracy and Trace Verifiability in Mathematical Reasoning},\n  author        = {Bei, Sisong and Arbuzov, Mikhail L and Dong, Ziwei and Han, Yan and Kalaev, Dmitri and Zhang, Yanxin and Shvets, Alexey},\n  year          = {2026},\n  note          = {ICLR 2027 submission},\n  url           = {https://telegrapher.ai/research/answer-accuracy-and-trace-verifiability/}\n}",
  "gist": [
    {
      "label": "Claim",
      "text": "Math traces that passed a deterministic checker had the right answer about one time in three"
    },
    {
      "label": "TL;DR",
      "text": "On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy."
    },
    {
      "label": "Method",
      "text": "Outputs of Qwen3.5-4B checkpoints were scored twice, by a deterministic trace checker and by a separate reference-answer grader, in a panel of 25 trained checkpoints on a competition math suite at 2,048 tokens and in six seed pairs that trained the trace interface and free-form chain-of-thought with the same outcome-only reward, evaluated on 200 held-out GSM8K and MATH problems at a 1,024-token cap."
    },
    {
      "label": "Key result",
      "text": "Accuracy among checker-accepted outputs: 35.0%; Share of correct answers whose trace was accepted: 21.8%; Held-out accuracy change from adopting the trace interface: −15.6 pp"
    },
    {
      "label": "Why it matters",
      "text": "Whoever filters or reports model outputs by a checker's verdict should read the joint table, not the verified fraction."
    },
    {
      "label": "Limits",
      "text": "The trained-model evidence covers one model, Qwen3.5-4B, on math, and acceptance means passing one checker's supported grammar and arithmetic rules; a checked claim is a specified test on a declared statement, not a check that the answer follows from the givens."
    },
    {
      "label": "Status",
      "text": "ICLR 2027 submission, August 2026"
    },
    {
      "label": "Read",
      "text": "reader /research/answer-accuracy-and-trace-verifiability/read/, PDF /papers/answer-accuracy-and-trace-verifiability/answer-accuracy-and-trace-verifiability.pdf"
    }
  ],
  "note": {
    "slug": "answer-accuracy-and-trace-verifiability",
    "claim_title": "Math traces that passed a deterministic checker had the right answer about one time in three",
    "meta_description": "A deterministic checker for a small model's math traces: accepted outputs were correct 35.0% of the time, and switching to the trace format cost accuracy.",
    "tldr": "On a small model's math reasoning, traces a deterministic checker accepted had the right answer about one time in three, and training the model to write checkable traces raised acceptance while lowering accuracy.",
    "gist": "A reasoning trace that a program can check gives a reader a second signal besides the answer, but passing the check and being right are separate outcomes. The study trained Qwen3.5-4B to write a compact trace language whose equations and checks a deterministic checker evaluates, and graded the final answers of the same outputs against references. Across 17,025 pooled outputs from 25 trained checkpoints, accepted outputs were correct 35.0% of the time against 18.38% overall. That is a real enrichment; most accepted answers were still wrong, and the filter kept about one correct answer in five. In six seed-paired runs against free-form chain-of-thought under the same outcome-only reward, the trace interface raised acceptance by about 34 points and lowered held-out accuracy by about 16 at matched training examples, with the same signs in all six pairs and already at the first measured checkpoint. A pre-registered negative association between accuracy and acceptance across checkpoints did not confirm. The evidence covers one 4B model on math, and the interface comparison changes language, prompt, constraints and answer channel together.",
    "method": "Outputs of Qwen3.5-4B checkpoints were scored twice, by a deterministic trace checker and by a separate reference-answer grader, in a panel of 25 trained checkpoints on a competition math suite at 2,048 tokens and in six seed pairs that trained the trace interface and free-form chain-of-thought with the same outcome-only reward, evaluated on 200 held-out GSM8K and MATH problems at a 1,024-token cap.",
    "summary_html": [
      "One trace spends 29 prose steps hunting for the value of a sum, tries <code>80? 81?</code> and settles on 80. Then it writes <code>EQ[eq_a]: S = 80</code> and <code>CHECK[c1]: arith: eq_a.value == 80</code>. Both claims check, the revised checker returns VERIFIED, and the grader marks the answer wrong, because the check confirms agreement with a value the trace asserted, not how that value follows from the givens. The trace language makes part of a model's output executable: <code>GIVEN</code> binds variables, <code>GOAL</code> names the target, <code>STEP</code> carries prose, <code>EQ</code> and <code>CHECK</code> expose claims a program evaluates, and <code>ANS</code> gives the answer. A deterministic checker's accepting verdict is called VERIFIED, and its share of all outputs the <em>verified fraction</em>. A separate grader takes another route through the same output: it extracts the final answer and compares it with the reference. Both were run over Qwen3.5-4B outputs in two studies. A panel of 25 trained checkpoints at 2,048 tokens, scored with the original v1 checker, supplies the joint distribution. Six seed pairs train the trace interface and free-form chain-of-thought from the same base model, data order and outcome-only reward, with no reward for acceptance.",
      "On the panel's 17,025 pooled outputs, acceptance picked out a more accurate subset: 35.0% of accepted outputs were correct, against 18.38% of all outputs. A real filter, and a leaky one. Wrong-and-accepted outputs outnumbered correct-and-accepted ones, and the rejected column held 2,446 correct answers, so the filter kept 21.8% of the correct answers available. Adopting the trace interface changed the pool itself. Matched on training examples, it raised verified fraction by 34.1 points and lowered held-out accuracy by 15.6; matched on training tokens, the gaps were 35.1 and 19.1, and all six pairs shared both signs under both matches. Since free-form output does not target the trace language, its acceptance is zero and the acceptance gain equals the trace arm's own rate. Both gaps were already present at the first measured checkpoint. Within every trace run accuracy ended higher than it started, while acceptance moved up, down or not at all. Across checkpoints, the registered test for a negative accuracy–acceptance association did not confirm: the rank correlation was −0.087 against a target of −0.40 or lower, and none of 12 checkpoint pairs matched on accuracy and output length differed by the required five points of verified fraction."
    ],
    "key_numbers": [
      {
        "label": "Accuracy among checker-accepted outputs",
        "value": "35.0%",
        "context": "v1 checker, 17,025 pooled checkpoint–problem outputs at 2,048 tokens; 18.38% across all outputs"
      },
      {
        "label": "Share of correct answers whose trace was accepted",
        "value": "21.8%",
        "context": "same panel; the rejected outputs held 2,446 correct answers",
        "bad": true
      },
      {
        "label": "Held-out accuracy change from adopting the trace interface",
        "value": "−15.6 pp",
        "context": "mean of six seed pairs at matched training examples, 200 held-out problems; −19.1 pp at matched training tokens",
        "bad": true
      },
      {
        "label": "Verified-fraction change from adopting the trace interface",
        "value": "+34.1 pp",
        "context": "same six pairs and match; free-form acceptance is zero, so this is the trace arm's own rate; +35.1 pp at matched training tokens"
      },
      {
        "label": "Rank correlation of checkpoint accuracy with acceptance",
        "value": "−0.087",
        "context": "25 checkpoints; 95% interval −0.4576 to +0.3091; the registered target was −0.40 or lower with an interval excluding zero"
      }
    ],
    "editorial_html": [
      "Whoever filters or reports model outputs by a checker's verdict should read the joint table, not the verified fraction. A verified fraction says how many traces passed. It does not say how many answers were right, and on this panel the table puts two costs side by side: the error left among kept answers, and the correct answers thrown away. There is also a quieter problem. When the checker evaluates statements the model itself wrote, an accepting verdict can rest on self-assertion, as in the trace that bound its guessed answer to an equation and then checked the equation against the guess.",
      "Choosing a reasoning format for its auditability is the second decision the results bear on. Under the same outcome-only reward, the format that made outputs checkable also produced fewer correct answers, and the accuracy gap persisted across the recorded checkpoints even as trace accuracy improved. The checkpoint result points the other way. Among models already writing traces, the data do not establish that more accurate checkpoints pass the checker less often, so the interface cost is not evidence of a general trade-off between accuracy and checkability.",
      "<em>Telegraph English</em> rewrites a model's input context into compact symbolic statements; this interface structures the reasoning a model generates instead. <em>Telegraph Reasoning</em> measured how many injected errors a rule-based linter catches in reasoning traces, which is a different question from what a checker's verdict says about a model's own answers. And <em>Why Additive Process Rewards Wash Out in Group-Normalized Reinforcement Learning</em> finds trajectory-additive process rewards suppressed under group normalization; the paired runs here kept the reward on outcomes alone and measured acceptance beside accuracy."
    ],
    "limitations": "The trained-model evidence covers one model, Qwen3.5-4B, on math, and acceptance means passing one checker's supported grammar and arithmetic rules; a checked claim is a specified test on a declared statement, not a check that the answer follows from the givens. The interface comparison changes the trace language, prompting, decoding constraints and answer channel together, so it cannot say which component costs accuracy. Its first measurement comes at step 100, after training had begun. The study grew from three seed pairs to six after the first three results were visible, the paired means carry no population-level interval, and the trajectories are repeated readings of those six pairs rather than further replicates. The paired export does not record its checker version, and missing raw paired outputs and trainer state limit replay. The checkpoint test deviated from its registration too: it resampled checkpoints that share training lineages instead of problems, and matched pairs greedily rather than optimally.",
    "concept_terms": [
      "reasoning trace verification",
      "deterministic checker",
      "verified fraction",
      "answer accuracy",
      "outcome-only reinforcement learning",
      "pre-registration"
    ],
    "references": [
      "Wei et al. (2022). Chain-of-thought prompting elicits reasoning in large language models.",
      "Lightman et al. (2024). Let’s verify step by step.",
      "Cobbe et al. (2021). Training verifiers to solve math word problems.",
      "Gao et al. (2023). PAL: Program-aided language models.",
      "Shao et al. (2024). DeepSeekMath: Pushing the limits of mathematical reasoning in open language models."
    ],
    "source_used": "local_pdf_text",
    "source_text_ref": "C:/Users/mikea/SCRIPTS/telegrapher-site/work/paper_text/answer-accuracy-and-trace-verifiability.txt",
    "verification": {
      "claims": [
        {
          "claim": "Checker-accepted outputs had a correct answer about one time in three (claim_title, tldr, gist)",
          "verdict": "CONFIRMED",
          "source_quote": "An output retained by this rule therefore has a correct answer about one time in three."
        },
        {
          "claim": "Both studies use Qwen3.5-4B, a small model, on math reasoning",
          "verdict": "CONFIRMED",
          "source_quote": "Both start from Qwen3.5-4B"
        },
        {
          "claim": "Trace language: GIVEN binds variables, GOAL names the target, STEP carries prose, EQ and CHECK expose claims, ANS gives the answer",
          "verdict": "CONFIRMED",
          "source_quote": "GIVEN binds variables, GOAL names the target, and ANS supplies the answer. STEP carries prose; EQ and CHECK expose explicit claims."
        },
        {
          "claim": "The checker's accepting verdict is called VERIFIED and its share of all outputs the verified fraction",
          "verdict": "CONFIRMED",
          "source_quote": "accepting verdict VERIFIED, and the share of all generated outputs receiving it the verified fraction"
        },
        {
          "claim": "A separate grader extracts the final answer and compares it with the reference",
          "verdict": "CONFIRMED",
          "source_quote": "The grader extracts the final answer and compares it with the problem"
        },
        {
          "claim": "Example trace has 29 prose steps",
          "verdict": "CONFIRMED",
          "source_quote": "29 prose steps are outside the claim inventory"
        },
        {
          "claim": "Example trace tries 80? 81? and settles on 80",
          "verdict": "CONFIRMED",
          "source_quote": "STEP[s16]: Try integer values. 80? 81?"
        },
        {
          "claim": "Example trace writes EQ[eq_a]: S = 80 and CHECK[c1]: arith: eq_a.value == 80",
          "verdict": "CONFIRMED",
          "source_quote": "EQ[eq_a]: S = 80 CHECK[c1]: arith: eq_a.value == 80"
        },
        {
          "claim": "The revised checker (v1.1) returned VERIFIED on that trace and the grader marked the answer wrong",
          "verdict": "CONFIRMED",
          "source_quote": "was graded wrong. Its v1 verdict was invalid; v1.1 returned VERIFIED."
        },
        {
          "claim": "The check confirms agreement with an asserted value, not how that value follows from the givens",
          "verdict": "CONFIRMED",
          "source_quote": "Their agreement explains acceptance without establishing how that value follows from the givens."
        },
        {
          "claim": "An accepting verdict can rest on self-assertion, as in the S = 80 trace",
          "verdict": "CONFIRMED",
          "source_quote": "This is the self-assertion case illustrated in Figure 1."
        },
        {
          "claim": "Panel of 25 trained checkpoints at 2,048 tokens on a competition math suite",
          "verdict": "CONFIRMED",
          "source_quote": "25 non-pilot checkpoints Competition suite; primary AMC and Olympiad subset 2,048 tokens"
        },
        {
          "claim": "The panel is scored with the original v1 checker and supplies the joint distribution",
          "verdict": "CONFIRMED",
          "source_quote": "The panel supplies the joint answer–acceptance distribution"
        },
        {
          "claim": "17,025 pooled checkpoint-problem outputs",
          "verdict": "CONFIRMED",
          "source_quote": "17,025 pooled checkpoint–problem outputs at the 2,048-token budget"
        },
        {
          "claim": "35.0% of accepted outputs were correct against 18.38% of all outputs",
          "verdict": "CONFIRMED",
          "source_quote": "35.0% of accepted outputs have correct answers, compared with 18.38% of the full pool"
        },
        {
          "claim": "Wrong-and-accepted outputs outnumbered correct-and-accepted ones, so most accepted answers were wrong",
          "verdict": "CONFIRMED",
          "source_quote": "the wrong-and-accepted cell is larger than the correct-and-accepted cell"
        },
        {
          "claim": "The rejected column held 2,446 correct answers",
          "verdict": "CONFIRMED",
          "source_quote": "The rejected column contains 2,446 correct answers"
        },
        {
          "claim": "The filter kept 21.8% of correct answers, about one in five",
          "verdict": "CONFIRMED",
          "source_quote": "the filter retains only 21.8% of correct answers"
        },
        {
          "claim": "Acceptance is a real enrichment of the retained pool",
          "verdict": "CONFIRMED",
          "source_quote": "acceptance enriches the retained pool while rejecting most correct answers"
        },
        {
          "claim": "Six seed pairs train the trace interface and free-form chain-of-thought with the same outcome-only reward and no reward for acceptance",
          "verdict": "CONFIRMED",
          "source_quote": "Both arms use outcome-only group-relative policy optimization (Shao et al., 2024; DeepSeek-AI et al., 2025), with no reward for checker acceptance."
        },
        {
          "claim": "Paired runs share base model and data order",
          "verdict": "CONFIRMED",
          "source_quote": "The paired study holds the base model, data order, paired seeds, reward and generation allowance fixed"
        },
        {
          "claim": "Paired runs evaluated on 200 held-out GSM8K and MATH problems",
          "verdict": "CONFIRMED",
          "source_quote": "split into 813 training and 200 held-out problems shared by both arms"
        },
        {
          "claim": "Paired evaluation capped at 1,024 tokens",
          "verdict": "CONFIRMED",
          "source_quote": "Generation remains capped at 1,024 tokens."
        },
        {
          "claim": "At matched training examples: verified fraction +34.1 pp and held-out accuracy -15.6 pp (gist: about 34 and about 16)",
          "verdict": "CONFIRMED",
          "source_quote": "lowers accuracy by 15.6 percentage points and raises verified fraction by 34.1 points on average"
        },
        {
          "claim": "At matched training tokens: accuracy -19.1 pp and verified fraction +35.1 pp",
          "verdict": "CONFIRMED",
          "source_quote": "the mean accuracy difference is −19.1 points and the verified-fraction difference is +35.1 points"
        },
        {
          "claim": "All six pairs share both signs under both matches",
          "verdict": "CONFIRMED",
          "source_quote": "All six directions agree for each endpoint under each normalization."
        },
        {
          "claim": "Free-form acceptance is zero, so the acceptance gain equals the trace arm's own rate",
          "verdict": "CONFIRMED",
          "source_quote": "Free-form acceptance is always zero, so acceptance differences equal the trace-arm rates."
        },
        {
          "claim": "Both gaps were already present at the first measured checkpoint",
          "verdict": "CONFIRMED",
          "source_quote": "At the first measured checkpoint, every pair already has higher trace acceptance and lower trace accuracy than free-form reasoning."
        },
        {
          "claim": "The accuracy gap persisted across the recorded checkpoints even as trace accuracy improved (revised wording; Table 6 has trace accuracy below free-form at all 15 checkpoints in all six pairs)",
          "verdict": "CONFIRMED",
          "source_quote": "Those between-interface directions persist across the recorded checkpoints"
        },
        {
          "claim": "Within every trace run accuracy ended higher; acceptance moved up, down or not at all",
          "verdict": "CONFIRMED",
          "source_quote": "Within the trace arm, every run ends with higher answer accuracy than at its first measurement. Acceptance changes upward, downward or remains unchanged across those runs."
        },
        {
          "claim": "The pre-registered negative accuracy-acceptance association across checkpoints did not confirm",
          "verdict": "CONFIRMED",
          "source_quote": "the accuracy–checkability trade-off we pre-registered did not confirm"
        },
        {
          "claim": "Rank correlation -0.087 over 25 checkpoints, 95% interval -0.4576 to +0.3091",
          "verdict": "CONFIRMED",
          "source_quote": "The observed correlation is −0.087, with a 95% interval from −0.4576 to +0.3091."
        },
        {
          "claim": "Registered target: rank correlation -0.40 or lower with an interval excluding zero",
          "verdict": "CONFIRMED",
          "source_quote": "Confirmation required a rank correlation at or below −0.40 with an interval excluding zero"
        },
        {
          "claim": "None of 12 matched checkpoint pairs reached the required five points of verified fraction",
          "verdict": "CONFIRMED",
          "source_quote": "None of the 12 saved pairs reaches that margin even in absolute magnitude"
        },
        {
          "claim": "Checkpoint pairs were matched on accuracy and output length (revised from 'accuracy-matched')",
          "verdict": "CONFIRMED",
          "source_quote": "Eligible pairs differ by at most 2.0 accuracy points and 10% in emitted bytes."
        },
        {
          "claim": "The joint table puts two costs side by side: error among kept answers and correct answers discarded",
          "verdict": "CONFIRMED",
          "source_quote": "the joint table supplies two operational costs: error among retained answers and correct answers discarded by the filter"
        },
        {
          "claim": "The interface cost is not evidence of a general accuracy-checkability trade-off",
          "verdict": "CONFIRMED",
          "source_quote": "The observed interface contrast can support a comparison of generation procedures without establishing a general accuracy–acceptance trade-off across the panel."
        },
        {
          "claim": "Telegraph English rewrites input context into compact symbolic statements; this interface structures generated reasoning (this paper's related work; consistent with the telegraph-english abstract)",
          "verdict": "CONFIRMED",
          "source_quote": "Telegraph English rewrites context into compact symbolic statements"
        },
        {
          "claim": "Telegraph Reasoning measured how many injected errors a rule-based linter catches in reasoning traces (checked against the telegraph-reasoning abstract in abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "On a corpus of 266 traces with injected errors, the linter catches 99.5 percent of the errors."
        },
        {
          "claim": "Why Additive Process Rewards Wash Out finds trajectory-additive process rewards suppressed under group normalization (checked against the additive-process-rewards abstract in abstracts.yaml)",
          "verdict": "CONFIRMED",
          "source_quote": "trajectory-additive process rewards are suppressed, in two regimes"
        },
        {
          "claim": "The paired runs kept the reward on outcomes alone (revised from 'the runs here': the panel's a_traj and shuffled reward definitions are unspecified)",
          "verdict": "CONFIRMED",
          "source_quote": "Both arms use outcome-only group-relative policy optimization"
        },
        {
          "claim": "The trained-model evidence covers one model, Qwen3.5-4B, on math",
          "verdict": "CONFIRMED",
          "source_quote": "The trained-model evidence covers one model family at one scale"
        },
        {
          "claim": "Acceptance means passing one checker's supported grammar and arithmetic rules",
          "verdict": "CONFIRMED",
          "source_quote": "supported grammar and arithmetic rules"
        },
        {
          "claim": "A checked claim is a specified test on a declared statement, not a check that the answer follows from the givens",
          "verdict": "CONFIRMED",
          "source_quote": "a checked classification refers to a specified implemented test on a declared statement"
        },
        {
          "claim": "The interface comparison changes trace language, prompting, decoding constraints and answer channel together, so it cannot isolate which component costs accuracy",
          "verdict": "CONFIRMED",
          "source_quote": "separating the contributions of its prompt, constraints and answer channel requires component comparisons"
        },
        {
          "claim": "The first measurement comes at step 100, after training had begun",
          "verdict": "CONFIRMED",
          "source_quote": "The first observation follows training start"
        },
        {
          "claim": "The study grew from three to six seed pairs after the first three results were visible",
          "verdict": "CONFIRMED",
          "source_quote": "The paired study expanded to six seed pairs after the first three pairs were visible."
        },
        {
          "claim": "The paired means carry no population-level interval",
          "verdict": "CONFIRMED",
          "source_quote": "No population-level interval is supplied for these paired means."
        },
        {
          "claim": "The trajectories are repeated readings of the six pairs, not further replicates",
          "verdict": "CONFIRMED",
          "source_quote": "The trajectories are repeated measurements within the six pairs, rather than additional independent training replicates."
        },
        {
          "claim": "The paired export does not record its checker version; missing raw paired outputs and trainer state limit replay",
          "verdict": "CONFIRMED",
          "source_quote": "missing raw paired outputs and trainer state limit replay of the paired study"
        },
        {
          "claim": "The checkpoint test resampled checkpoints sharing training lineages instead of problems and matched greedily rather than optimally",
          "verdict": "CONFIRMED",
          "source_quote": "The implemented analysis resamples checkpoints with shared training lineages instead of the registered problems, and uses greedy rather than optimal matching."
        },
        {
          "claim": "All five references (Wei et al. 2022, Lightman et al. 2024, Cobbe et al. 2021, Gao et al. 2023, Shao et al. 2024) appear in the paper's reference list",
          "verdict": "CONFIRMED",
          "source_quote": "Chain-of-thought prompting elicits reasoning in large language models."
        }
      ],
      "revised": true,
      "status": "pass"
    }
  }
}