Question review agent

ONC2-7195 ยท Replaces the report fixer (fix-question-task-v2) with one pipeline for every course and format. All numbers from prod data, read-only, on 2026-09-28.

79.7%benchmark cases handled correctly (363 cases)
1.8%scored harmful; by hand, 1 likely miss and 2 judgment calls
0wrong keys applied (8 wrong proposals stopped by the checks)
$0.033per review, median 24s

What was wrong

How it works now

  1. Load the whole question. The reported row plus its set (case study, TBS, question set, reading set), hidden parts dropped, images checked to be reachable.
  2. Code checks the structure. Per-format checks, and the stored key is submitted to the real grader and must score 100%. Code findings are facts the review must resolve.
  3. Blind solve. GPT-6.1-sol answers every keyed part with the key, explanation and report withheld. (The approach the September explanation backfill used for key disputes.)
  4. Review agent. An AI SDK ToolLoopAgent on GPT-6.1-sol via AI Gateway (each call capped at 8,000 output tokens, retried and finally sent to GPT-6-sol if it returns nothing parseable), with a calculator and Perplexity search, judges the report against the key, the blind answer and the course's exam conventions, and returns a typed verdict: no change, fix, disable or escalate.
  5. Deterministic, checked writes. Fixes are exact text replacements or a new key in the format's own terms, written into every copy (content, option rows, flattened text, explanation copies). A fix goes live only if: a key changes only when the blind solve of the original question (at 0.6+ confidence) or a structural check disagreed with it; every changed key is reached by two blind solvers from different model families (GPT-6.1-sol and Claude Opus 5.5) at 0.8+ confidence; every part whose text changed is re-solved from the new text and still reaches its key; the question lints clean; and the rebuilt explanation passes the house-format checks. Disabling needs a structural finding or a confident blind-solve disagreement on the item the reviewer blames. Otherwise a person decides.
  6. Record and notify. One transaction that refuses a question edited since it was read, never moves a disabled question back into practice, and keeps the previous fields and status in metadata.question_review so a run can be reverted; then audit rows, a Slack summary and the learner email.

Vercel's eve was considered and not used: it is a separate agent server with its own workflow store, sandbox and auth, which would duplicate Trigger.dev for a job that is one bounded review per report.

Benchmark: 363 frozen cases

Built read-only from prod (scripts/question-review/build-benchmark.ts, data in scripts/question-review/data/benchmark.jsonl.gz). Scored per case as correct, to a person (held or escalated when a fix or no change was expected) or harmful (a wrong key applied, a sound question disabled, or a known defect left live).

Case typeCasesCorrectTo a personHarmful
Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves)3781%19%0
Real key errors from the September backfill, restored2850%46%1
Keys the backfill disputed, then verified (learner argues the wrong answer)3577%23%0
Sound question, mistaken report5074%24%1
Explanation from another question3491%9%0
Stem cut off3591%0%3
Key moved to a distractor7081%19%0
Option rows disagree with the content key1486%14%0
Image stripped from an image question19100%0%0
Image present, "image missing" report1662%31%1
Linked image is dead in storage2100%0%0

Every harmful case, read by hand

The scorer calls a case harmful when the final key differs from the label, a sound question is disabled, or a known defect stays live. Each one was checked against the question itself:

CaseCourse / formatOutcomeAdjudication
Real key errors from the September backfill, restoredLSAT / mcq_singlefixedDiffers from the label, not harmful. Both B and D follow from the premises, so the question had two right answers. The label's fix rewrites D and keys B; the review kept D and rewrote B into an unsupported "every" claim, so one answer is right.
Image present, "image missing" reportCFA / case_studydisabledReal defect the label missed. The figure's one- and two-year spot rates give a forward rate near 4.00%, but the key is 4.21%, so the case was rightly taken out of practice.
Stem cut offCFA / matrix_gridno_changeWeak seed. The cut left the matching item answerable from its rows; both solvers answered it with full confidence.
Sound question, mistaken reportCFA / fill_blankfixedNot harmful. "USD 3.0 million" became 3 with the unit added to the stem; the old key could not be entered on the numeric keypad.
Stem cut offIndian Medical PG / mcq_singleno_changeReal miss. The stem stops mid-sentence ("The blood-testis barrier is"), and the reviewer called it answerable from the choices. The GPT-6-sol run rewrote it.
Stem cut offNCLEX-RN / ordered_responseno_changeReal miss, low harm. The ordering instruction is cut off, but the format itself asks the learner to order the steps. The GPT-6-sol run restored it.

Real reports: old agent vs new

Labelled real reports only (both independent solvers agree). These are Indian PG MCQs, the old agent's home ground; on every other course and format the old agent did nothing.

Stored keyCasesOld agentNew agent
Right (both solvers agree with it)18left alone 18, changed wrongly 0left alone 18, changed wrongly 0
Wrong (both solvers agree on another answer)19fixed 12, not fixed 7fixed 12, disabled 4, to a person 3, missed 0

Models on AI Gateway

The same pipeline with each model doing the blind solve and the review, on a stratified 120-case subset (12 per case type). Key changes are confirmed by Claude Opus 5.5 in every run. The three alternatives ran on an earlier build of the checks, so treat the comparison as rough.

Review + solve modelScored casesCorrectTo a personHarmfulKey changes against the labelCost / reviewMedian time
GPT-6.1-sol (default)11979%18%31$0.03122s
GPT-6-sol (fallback)11982%15%31$0.03425s
Claude Opus 5.512078%18%52$0.12832s
Gemini 3.8 Flash12069%28%31$0.03534s
GPT-6-luna12078%18%51$0.00621s

By course and format

Course / formatCasesCorrectTo a personHarmful
Bar Exam / cloze_dropdown978%22%0
Bar Exam / integrated_question_set250%50%0
Bar Exam / matrix_grid580%20%0
Bar Exam / mcq_single1493%7%0
Bar Exam / ordered_response5100%0%0
Bar Exam / performance_task20%100%0
Bar Exam / sata5100%0%0
Bar Exam / short_answer1100%0%0
CFA / case_study667%17%1
CFA / fill_blank956%33%1
CFA / highlight_hot_spot4100%0%0
CFA / matrix_grid560%20%1
CFA / mcq_single4171%29%0
CFA / ordered_response580%20%0
CFA / sata11100%0%0
CPA / cloze_dropdown786%14%0
CPA / cpa_tbs250%50%0
CPA / data_entry_grid771%29%0
CPA / mcq_single1486%14%0
Indian Medical PG / mcq_single4680%17%1
LSAT / argumentative_essay10%100%0
LSAT / mcq_single2095%0%1
LSAT / reading_comprehension2100%0%0
MCAT / mcq_single1060%40%0
NCLEX-RN / case_study667%33%0
NCLEX-RN / cloze_dropdown1090%10%0
NCLEX-RN / fill_blank9100%0%0
NCLEX-RN / highlight_hot_spot1090%10%0
NCLEX-RN / image_hot_spot5100%0%0
NCLEX-RN / matrix_grid1090%10%0
NCLEX-RN / mcq_single1669%31%0
NCLEX-RN / ordered_response1080%10%1
NCLEX-RN / sata1656%44%0
UK Medical PG / mcq_single5100%0%0
US Medical PG / mcq_single1090%10%0

Every report since 2024

2650 reports (1843 from 480 learners, 807 from the system user) on 2065 questions; 381 questions were reported more than once. 817 learner reports carry a written note, 325 a screenshot. Jev classified every note by what the learner claims (200 name a specific answer; 7 cite a source):

What the learner claims (Jev)NotesShare
The key is wrong / another answer is right29536%
Image missing, wrong or unreadable17121%
A choice is wrong, duplicated or missing9211%
Stem unclear, incomplete or has a typo9011%
Note does not say what is wrong8010%
Explanation wrong or mismatched567%
App problem, not the question192%
Outdated or off-syllabus142%
Old agent action within 3 hoursReportsShare
no change97937%
key changed56821%
explanation only55921%
stem edited35613%
image changed1445%
disabled442%

Not in this PR: validating the whole bank

This PR reviews one question at a time, when a learner reports it. A pass over all 113K live question sets is a separate PR: a targeted set of Jev checks per course and question type, run across the bank, with the full review reserved for what Jev flags. Jev (TypeSafe's System One model) told wrong keys from right ones with AUC 0.86 in a first measurement on this benchmark, without reasoning, at about $0.04 per million input tokens.

What else this found