Question review agent

ONC2-7195 ยท Replaces the report fixer (fix-question-task-v2) with one pipeline for every course and format. All numbers from prod data, read-only, on 2026-09-28.

80%benchmark cases handled correctly (363 cases)
2.6%scored harmful; by hand, 1 likely miss and 2 judgment calls
0wrong keys applied (9 wrong proposals stopped by the checks)
$0.035per review, median 26s

What was wrong

How it works now

  1. Load the whole question. The reported row plus its set (case study, TBS, question set, reading set), hidden parts dropped, images checked to be reachable.
  2. Code checks the structure. Per-format checks, and the stored key is submitted to the real grader and must score 100%. Code findings are facts the review must resolve.
  3. Blind solve. GPT-6-sol answers every keyed part with the key, explanation and report withheld. (The approach the September explanation backfill used for key disputes.)
  4. Review agent. An AI SDK ToolLoopAgent on GPT-6-sol via AI Gateway, with a calculator and Perplexity search, judges the report against the key, the blind answer and the course's exam conventions, and returns a typed verdict: no change, fix, disable or escalate.
  5. Deterministic, checked writes. Fixes are exact text replacements or a new key in the format's own terms, written into every copy (content, option rows, flattened text, explanation copies). A fix goes live only if: a key changes only when the blind solve of the original question (at 0.6+ confidence) or a structural check disagreed with it; every changed key is reached by two blind solvers from different model families (GPT-6-sol and Claude Opus 5.5) at 0.8+ confidence; every part whose text changed is re-solved from the new text and still reaches its key; the question lints clean; and the rebuilt explanation passes the house-format checks. Disabling needs a structural finding or a confident blind-solve disagreement on the item the reviewer blames. Otherwise a person decides.
  6. Record and notify. One transaction that refuses a question edited since it was read, never moves a disabled question back into practice, and keeps the previous fields and status in metadata.question_review so a run can be reverted; then audit rows, a Slack summary and the learner email.

Vercel's eve was considered and not used: it is a separate agent server with its own workflow store, sandbox and auth, which would duplicate Trigger.dev for a job that is one bounded review per report.

Benchmark: 363 frozen cases

Built read-only from prod (scripts/question-review/build-benchmark.ts, data in scripts/question-review/data/benchmark.jsonl.gz). Scored per case as correct, to a person (held or escalated when a fix or no change was expected) or harmful (a wrong key applied, a sound question disabled, or a known defect left live).

Case typeCasesCorrectTo a personHarmful
Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves)3778%11%4
Real key errors from the September backfill, restored2846%54%0
Keys the backfill disputed, then verified (learner argues the wrong answer)3571%23%2
Sound question, mistaken report5080%18%1
Key moved to a distractor7081%19%0
Explanation from another question3491%9%0
Stem cut off3597%0%1
Option rows disagree with the content key1486%14%0
Image stripped from an image question19100%0%0
Image present, "image missing" report1662%31%1
Linked image is dead in storage2100%0%0

Every harmful case, read by hand

The scorer calls a case harmful when the final key differs from the label, a sound question is disabled, or a known defect stays live. Each one was checked against the question itself:

CaseCourse / formatOutcomeAdjudication
Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves)Indian Medical PG / mcq_singlefixedLikely a real miss. Indian PG texts describe flat vegetations in valve pockets as Libman-Sacks, as both labelling models answered; the review's own solver agreed with the stored NBTE key, so the review kept it and rebuilt the explanation. Needs a person.
Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves)Indian Medical PG / mcq_singlefixedFine. The labelled answer was a second correct choice; the review replaced it with a true restrictive operation and tightened the stem, and both solvers reach the key.
Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves)Indian Medical PG / mcq_singlefixedJudgment call. Correct now, but the review narrowed the stem to keep the disputed key rather than changing the key.
Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves)Indian Medical PG / mcq_singlefixedJudgment call. Correct now, but the review narrowed the stem to keep the disputed key rather than changing the key.
Keys the backfill disputed, then verified (learner argues the wrong answer)CFA / satafixedNot harmful. The same three statements stay keyed; one was reworded to be accurate, which the scorer counts as a key change.
Keys the backfill disputed, then verified (learner argues the wrong answer)CFA / mcq_singlefixedNot harmful. It fixed a distractor that was a second correct answer, and turned a sibling part's key "USD 8,920,000" into 8920000, the same value in a form learners can enter.
Stem cut offCFA / matrix_gridno_changeWeak seed. The cut left the matching item answerable from its rows; both solvers answered it with full confidence.
Sound question, mistaken reportCFA / fill_blankfixedNot harmful. "USD 3.0 million" became 3 with the unit added to the stem; the old key could not be entered on the numeric keypad.
Image present, "image missing" reportCFA / case_studydisabledReal defect the label missed. The figure's one- and two-year spot rates give a forward rate near 4.00%, but the key is 4.21%, so the case was rightly taken out of practice.

Real reports: old agent vs new

Labelled real reports only (both independent solvers agree). These are Indian PG MCQs, the old agent's home ground; on every other course and format the old agent did nothing.

Stored keyCasesOld agentNew agent
Right (both solvers agree with it)18left alone 18, changed wrongly 0left alone 18, changed wrongly 0
Wrong (both solvers agree on another answer)19fixed 12, not fixed 7fixed 11, disabled 2, to a person 2, missed 4

Models on AI Gateway

The same pipeline with each model doing the blind solve and the review, on a stratified 120-case subset (12 per case type). Key changes are confirmed by Claude Opus 5.5 in every run. The three alternatives ran on an earlier build of the checks, so treat the comparison as rough.

Review + solve modelScored casesCorrectTo a personHarmfulKey changes against the labelCost / reviewMedian time
GPT-6-sol (default)11982%15%31$0.03425s
Claude Opus 5.512078%18%52$0.12832s
Gemini 3.8 Flash12069%28%31$0.03534s
GPT-6-luna12078%18%51$0.00621s

By course and format

Course / formatCasesCorrectTo a personHarmful
Bar Exam / cloze_dropdown978%22%0
Bar Exam / integrated_question_set2100%0%0
Bar Exam / matrix_grid5100%0%0
Bar Exam / mcq_single1493%7%0
Bar Exam / ordered_response5100%0%0
Bar Exam / performance_task250%50%0
Bar Exam / sata580%20%0
Bar Exam / short_answer1100%0%0
CFA / case_study667%17%1
CFA / fill_blank967%22%1
CFA / highlight_hot_spot4100%0%0
CFA / matrix_grid560%20%1
CFA / mcq_single4168%29%1
CFA / ordered_response580%20%0
CFA / sata1182%9%1
CPA / cloze_dropdown786%14%0
CPA / cpa_tbs250%50%0
CPA / data_entry_grid757%43%0
CPA / mcq_single1493%7%0
Indian Medical PG / mcq_single4680%11%4
LSAT / argumentative_essay1100%0%0
LSAT / mcq_single2095%5%0
LSAT / reading_comprehension2100%0%0
MCAT / mcq_single1070%30%0
NCLEX-RN / case_study667%33%0
NCLEX-RN / cloze_dropdown1090%10%0
NCLEX-RN / fill_blank989%11%0
NCLEX-RN / highlight_hot_spot1090%10%0
NCLEX-RN / image_hot_spot5100%0%0
NCLEX-RN / matrix_grid1090%10%0
NCLEX-RN / mcq_single1675%25%0
NCLEX-RN / ordered_response1080%20%0
NCLEX-RN / sata1662%38%0
UK Medical PG / mcq_single580%20%0
US Medical PG / mcq_single1080%20%0

Screening every question

A full review costs about $0.035, so the sweep screens first and reviews only what is flagged. Measured on the benchmark's key errors (131) and sound keys (187):

ScreenKey errors caughtSound keys flagged
gpt-6-luna blind solve85.5%6.4%
Jev >= 0.375.6%7%
gpt-6-luna blind solve or Jev >= 0.387%11.2%
gpt-6-luna blind solve or Jev >= 0.585.5%7.5%
gpt-6-sol blind solve88.5%3.7%

GPT-6-luna costs about $0.00023 per question. Its misses are almost all key-copy mismatches, which the code checks catch, so lint plus the luna solve catch about 96% of key errors. Jev (TypeSafe's System One model) scores AUC 0.86 at telling wrong keys from right ones without reasoning; it misses calculation items, so it adds the non-key defect signal rather than being the key screen. Learners' answer data adds a third signal: 181 questions with 30+ answers where one distractor is picked more than twice as often as the key.

Every report since 2024

2650 reports (1843 from 480 learners, 807 from the system user) on 2065 questions; 381 questions were reported more than once. 817 learner reports carry a written note, 325 a screenshot. Jev classified every note by what the learner claims (200 name a specific answer; 7 cite a source):

What the learner claims (Jev)NotesShare
The key is wrong / another answer is right29536%
Image missing, wrong or unreadable17121%
A choice is wrong, duplicated or missing9211%
Stem unclear, incomplete or has a typo9011%
Note does not say what is wrong8010%
Explanation wrong or mismatched567%
App problem, not the question192%
Outdated or off-syllabus142%
Old agent action within 3 hoursReportsShare
no change97937%
key changed56821%
explanation only55921%
stem edited35613%
image changed1445%
disabled442%

Sweep of live questions (partial)

The sweep was stopped after CFA and most of CPA to save gateway credit; nothing from it has been written to prod. Its fixes were proposed while the checks were still being tightened, so the full sweep will review them again before anything is applied.

CourseLive setsScreenedPassed screenReviewed, soundFixedDisabledTo a person
CFA1808178915877393036
CPA168014651402391644

Proposed fixes: 90 key corrections (mostly CFA numeric keys stored with units, such as "USD 360,000" or "4.0x", that learners could not enter), 108 text corrections and 198 rebuilt explanations. Screen cost $1.0, review cost $12.19 for 3254 sets.

What else this found