Question review agent

ONC2-7195 ยท Replaces the report fixer (fix-question-task-v2) with one pipeline for every course and format. All numbers from prod data, read-only, on 2026-09-28.

80%benchmark cases handled correctly (363 cases)
0.6%scored harmful; by hand, 1 likely miss and 2 judgment calls
0wrong keys applied (10 wrong proposals stopped by the checks)
$0.042per review, median 29s

What was wrong

How it works now

  1. Load the whole question. The reported row plus its set (case study, TBS, question set, reading set), hidden parts dropped, images checked to be reachable.
  2. Code checks the structure. Per-format checks, and the stored key is submitted to the real grader and must score 100%. Code findings are facts the review must resolve.
  3. Blind solve. GPT-6.1-sol answers every keyed part with the key, explanation and report withheld. (The approach the September explanation backfill used for key disputes.)
  4. Review agent. An AI SDK ToolLoopAgent on GPT-6.1-sol via AI Gateway (each call capped at 8,000 output tokens, retried and finally sent to GPT-6-sol if it returns nothing parseable), with a calculator and Perplexity search, judges the report against the key, the blind answer and the course's exam conventions, and returns a typed verdict: no change, fix, disable or escalate.
  5. Deterministic, checked writes. Fixes are exact text replacements, a new key in the format's own terms, or a redrawn image (below), written into every copy (content, option rows, flattened text, explanation copies). A fix goes live only if: a key changes only when the blind solve of the original question (at 0.6+ confidence) or a structural check disagreed with it; every changed key is reached by two blind solvers from different model families (GPT-6.1-sol and Claude Opus 5.5) at 0.8+ confidence; every part whose text changed is re-solved from the new text and still reaches its key; the question lints clean; and the rebuilt explanation passes the house-format checks. Disabling needs a structural finding or a confident blind-solve disagreement on the item the reviewer blames. Otherwise a person decides.
  6. Record and notify. One transaction that refuses a question edited since it was read, never moves a disabled question back into practice, and keeps the previous fields and status in metadata.question_review so a run can be reverted; then audit rows, a Slack summary and the learner email.

Vercel's eve was considered and not used: it is a separate agent server with its own workflow store, sandbox and auth, which would duplicate Trigger.dev for a job that is one bounded review per report.

Image repair

About one learner note in five is about an image (171 of 817). The old fixer could search for, edit or draw a replacement. The new pipeline does it inside the same checks:

  1. The reviewer asks. It sees every image with its URL. If one is dead, wrong, unreadable or gives the answer away, it asks for a redrawn one and writes a complete brief: what it shows, and every label, unit and value a chart needs.
  2. gpt-image-2.5-flare draws it, another model family checks it. Flare is called through OpenAI directly (the AI Gateway does not serve it); Gemini 3 Pro Image through the gateway is the fallback. Claude Opus 5.5 checks the picture against the brief for missing or wrong labels, bars that do not match their numbers, unreadable text and anything that names the answer. A rejected image is drawn once more with the complaint.
  3. The question is solved again with the new image. The blind solve must still reach the key and no image on the question may be broken. Until then the image exists only in memory: nothing is stored or linked.
  4. Stored with the write. The image is stored, the row's linked asset and image_details point at it, and the previous ones are kept in the backup, so a revert restores them.

Not redrawn: image_hot_spot questions (their answer regions belong to the current image's pixels) and questions with no image URL at all (nothing says what belonged there). Those are still taken out of practice.

Three runs of the real pipeline on real questions

Each used a real benchmark question with a fault put in by hand, and the real models. Nothing was written.

QuestionFaultOutcomeTimeCost
CFA, mean-standard deviation diagramShown an unrelated chartFixed: the reviewer wrote a brief from the explanation, flare drew it, Claude accepted it, and the solve with it still reached the key40 s$0.10
MCAT, enzyme activity curveShown an unrelated chartSent to a person: the reviewer would not invent a curve whose data it could not recover15 s$0.02
CFA, the same diagramImage URL deadDisabled: the exact plotted values could not be recovered from the text13 s$0.01
The unrelated chart shown
Shown on the CFA question (wrong)
The redrawn diagram
Redrawn by the pipeline from the reviewer's brief
The question's original diagram
The question's original image, for comparison

In the benchmark

Both dead-image questions were redrawn and fixed, and all 19 questions whose image had been removed were still taken out of practice, since nothing says what belonged there. Of the 16 image questions labelled sound because the report was mistaken, 9 were left alone, 4 went to a person and 3 were redrawn. I looked at those three: one figure is titled "Long Straddle Expiration Profit Diagram" on a question asking which position Figure 1 shows, so it named the answer; one shows propylbenzene where the stem says ethylbenzene; the third has tranche labels the reviewer says do not match the question's classes, which I did not check myself. The benchmark counts them as edits to a sound question, but the first two are real defects the label missed.

The redrawn diagram is qualitative (positions, no numbers), which is all the explanation and the question support. The benchmark has only 2 dead-image cases, so its numbers say little about repair; the unit tests and these runs are the evidence.

Benchmark: 363 frozen cases

Built read-only from prod (scripts/question-review/build-benchmark.ts, data in scripts/question-review/data/benchmark.jsonl.gz). Scored per case as correct, to a person (held or escalated when a fix or no change was expected) or harmful (a wrong key applied, a sound question disabled, or a known defect left live).

Case typeCasesCorrectTo a personHarmful
Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves)3778%22%0
Real key errors from the September backfill, restored2857%43%0
Keys the backfill disputed, then verified (learner argues the wrong answer)3574%26%0
Explanation from another question3491%9%0
Key moved to a distractor7080%20%0
Sound question, mistaken report5076%22%1
Stem cut off3597%0%1
Option rows disagree with the content key1486%14%0
Image stripped from an image question19100%0%0
Image present, "image missing" report1656%44%0
Linked image is dead in storage2100%0%0

Every harmful case, read by hand

The scorer calls a case harmful when the final key differs from the label, a sound question is disabled, or a known defect stays live. Each one was checked against the question itself:

CaseCourse / formatOutcomeAdjudication
Sound question, mistaken reportCFA / fill_blankfixedNot harmful. "USD 3.0 million" became 3 with the unit added to the stem; the old key could not be entered on the numeric keypad.
Stem cut offCFA / matrix_gridno_changeWeak seed. The cut left the matching item answerable from its rows; both solvers answered it with full confidence.

Real reports: old agent vs new

Labelled real reports only (both independent solvers agree). These are Indian PG MCQs, the old agent's home ground; on every other course and format the old agent did nothing.

Stored keyCasesOld agentNew agent
Right (both solvers agree with it)18left alone 18, changed wrongly 0left alone 18, changed wrongly 0
Wrong (both solvers agree on another answer)19fixed 12, not fixed 7fixed 11, disabled 4, to a person 4, missed 0

Models on AI Gateway

The same pipeline with each model doing the blind solve and the review, on a stratified 120-case subset (12 per case type). Key changes are confirmed by Claude Opus 5.5 in every run. The three alternatives ran on an earlier build of the checks, so treat the comparison as rough.

Review + solve modelScored casesCorrectTo a personHarmfulKey changes against the labelCost / reviewMedian time
GPT-6.1-sol (default)11981%18%11$0.03927s
GPT-6-sol (fallback)11982%15%31$0.03425s
Claude Opus 5.512078%18%52$0.12832s
Gemini 3.8 Flash12069%28%31$0.03534s
GPT-6-luna12078%18%51$0.00621s

By course and format

Course / formatCasesCorrectTo a personHarmful
Bar Exam / cloze_dropdown967%33%0
Bar Exam / integrated_question_set250%50%0
Bar Exam / matrix_grid580%20%0
Bar Exam / mcq_single1493%7%0
Bar Exam / ordered_response5100%0%0
Bar Exam / performance_task20%100%0
Bar Exam / sata5100%0%0
Bar Exam / short_answer1100%0%0
CFA / case_study667%33%0
CFA / fill_blank967%22%1
CFA / highlight_hot_spot4100%0%0
CFA / matrix_grid560%20%1
CFA / mcq_single4173%27%0
CFA / ordered_response580%20%0
CFA / sata1191%9%0
CPA / cloze_dropdown771%29%0
CPA / cpa_tbs250%50%0
CPA / data_entry_grid771%29%0
CPA / mcq_single1493%7%0
Indian Medical PG / mcq_single4680%20%0
LSAT / argumentative_essay1100%0%0
LSAT / mcq_single20100%0%0
LSAT / reading_comprehension2100%0%0
MCAT / mcq_single1060%40%0
NCLEX-RN / case_study667%33%0
NCLEX-RN / cloze_dropdown1090%10%0
NCLEX-RN / fill_blank9100%0%0
NCLEX-RN / highlight_hot_spot1090%10%0
NCLEX-RN / image_hot_spot580%20%0
NCLEX-RN / matrix_grid1090%10%0
NCLEX-RN / mcq_single1669%31%0
NCLEX-RN / ordered_response1080%20%0
NCLEX-RN / sata1656%44%0
UK Medical PG / mcq_single5100%0%0
US Medical PG / mcq_single1090%10%0

Every report since 2024

2650 reports (1843 from 480 learners, 807 from the system user) on 2065 questions; 381 questions were reported more than once. 817 learner reports carry a written note, 325 a screenshot. Jev classified every note by what the learner claims (200 name a specific answer; 7 cite a source):

What the learner claims (Jev)NotesShare
The key is wrong / another answer is right29536%
Image missing, wrong or unreadable17121%
A choice is wrong, duplicated or missing9211%
Stem unclear, incomplete or has a typo9011%
Note does not say what is wrong8010%
Explanation wrong or mismatched567%
App problem, not the question192%
Outdated or off-syllabus142%
Old agent action within 3 hoursReportsShare
no change97937%
key changed56821%
explanation only55921%
stem edited35613%
image changed1445%
disabled442%

Not in this PR: validating the whole bank

This PR reviews one question at a time, when a learner reports it. A pass over all 113K live question sets is a separate PR: a targeted set of Jev checks per course and question type, run across the bank, with the full review reserved for what Jev flags. Jev (TypeSafe's System One model) told wrong keys from right ones with AUC 0.86 in a first measurement on this benchmark, without reasoning, at about $0.04 per million input tokens.

What else this found