Question review agent
ONC2-7195 ยท Replaces the report fixer (fix-question-task-v2) with one pipeline for every course and format. All numbers from prod data, read-only, on 2026-09-28.
What was wrong
- New formats never got reviewed. The fixer threw for every format except
mcq_singleand skipped every non-medical course, so reports on CFA, CPA, LSAT, Bar and NCLEX's new formats failed or were ignored. A report also sets the question toinvalid, and several practice queries only servevalidatedquestions, so a failed or skipped review took the question out of those pools for good. - It edited the wrong fields. New formats render and grade from
content; the old agent only editedquestion_text, the option rows and the explanation, so a key it "fixed" on a new-format question kept grading against the old key. - Its key changes were unchecked. It changed the key after 568 of 2650 reports (21%). 67 questions had their key changed after two separate reports, and in 30 of them the second fix undid the first.
- Old machinery. A Claude Agent SDK subprocess with six database-writing tools and a medical-only prompt on Sonnet 4.6; no benchmark.
How it works now
- Load the whole question. The reported row plus its set (case study, TBS, question set, reading set), hidden parts dropped, images checked to be reachable.
- Code checks the structure. Per-format checks, and the stored key is submitted to the real grader and must score 100%. Code findings are facts the review must resolve.
- Blind solve. GPT-6.1-sol answers every keyed part with the key, explanation and report withheld. (The approach the September explanation backfill used for key disputes.)
- Review agent. An AI SDK
ToolLoopAgenton GPT-6.1-sol via AI Gateway (each call capped at 8,000 output tokens, retried and finally sent to GPT-6-sol if it returns nothing parseable), with a calculator and Perplexity search, judges the report against the key, the blind answer and the course's exam conventions, and returns a typed verdict: no change, fix, disable or escalate. - Deterministic, checked writes. Fixes are exact text replacements, a new key in the format's own terms, or a redrawn image (below), written into every copy (content, option rows, flattened text, explanation copies). A fix goes live only if: a key changes only when the blind solve of the original question (at 0.6+ confidence) or a structural check disagreed with it; every changed key is reached by two blind solvers from different model families (GPT-6.1-sol and Claude Opus 5.5) at 0.8+ confidence; every part whose text changed is re-solved from the new text and still reaches its key; the question lints clean; and the rebuilt explanation passes the house-format checks. Disabling needs a structural finding or a confident blind-solve disagreement on the item the reviewer blames. Otherwise a person decides.
- Record and notify. One transaction that refuses a question edited since it was read, never moves a disabled question back into practice, and keeps the previous fields and status in
metadata.question_reviewso a run can be reverted; then audit rows, a Slack summary and the learner email.
Vercel's eve was considered and not used: it is a separate agent server with its own workflow store, sandbox and auth, which would duplicate Trigger.dev for a job that is one bounded review per report.
Image repair
About one learner note in five is about an image (171 of 817). The old fixer could search for, edit or draw a replacement. The new pipeline does it inside the same checks:
- The reviewer asks. It sees every image with its URL. If one is dead, wrong, unreadable or gives the answer away, it asks for a redrawn one and writes a complete brief: what it shows, and every label, unit and value a chart needs.
- gpt-image-2.5-flare draws it, another model family checks it. Flare is called through OpenAI directly (the AI Gateway does not serve it); Gemini 3 Pro Image through the gateway is the fallback. Claude Opus 5.5 checks the picture against the brief for missing or wrong labels, bars that do not match their numbers, unreadable text and anything that names the answer. A rejected image is drawn once more with the complaint.
- The question is solved again with the new image. The blind solve must still reach the key and no image on the question may be broken. Until then the image exists only in memory: nothing is stored or linked.
- Stored with the write. The image is stored, the row's linked asset and
image_detailspoint at it, and the previous ones are kept in the backup, so a revert restores them.
Not redrawn: image_hot_spot questions (their answer regions belong to the current image's pixels) and questions with no image URL at all (nothing says what belonged there). Those are still taken out of practice.
Three runs of the real pipeline on real questions
Each used a real benchmark question with a fault put in by hand, and the real models. Nothing was written.
| Question | Fault | Outcome | Time | Cost |
|---|---|---|---|---|
| CFA, mean-standard deviation diagram | Shown an unrelated chart | Fixed: the reviewer wrote a brief from the explanation, flare drew it, Claude accepted it, and the solve with it still reached the key | 40 s | $0.10 |
| MCAT, enzyme activity curve | Shown an unrelated chart | Sent to a person: the reviewer would not invent a curve whose data it could not recover | 15 s | $0.02 |
| CFA, the same diagram | Image URL dead | Disabled: the exact plotted values could not be recovered from the text | 13 s | $0.01 |



In the benchmark
Both dead-image questions were redrawn and fixed, and all 19 questions whose image had been removed were still taken out of practice, since nothing says what belonged there. Of the 16 image questions labelled sound because the report was mistaken, 9 were left alone, 4 went to a person and 3 were redrawn. I looked at those three: one figure is titled "Long Straddle Expiration Profit Diagram" on a question asking which position Figure 1 shows, so it named the answer; one shows propylbenzene where the stem says ethylbenzene; the third has tranche labels the reviewer says do not match the question's classes, which I did not check myself. The benchmark counts them as edits to a sound question, but the first two are real defects the label missed.
The redrawn diagram is qualitative (positions, no numbers), which is all the explanation and the question support. The benchmark has only 2 dead-image cases, so its numbers say little about repair; the unit tests and these runs are the evidence.
Benchmark: 363 frozen cases
Built read-only from prod (scripts/question-review/build-benchmark.ts, data in scripts/question-review/data/benchmark.jsonl.gz). Scored per case as correct, to a person (held or escalated when a fix or no change was expected) or harmful (a wrong key applied, a sound question disabled, or a known defect left live).
| Case type | Cases | Correct | To a person | Harmful |
|---|---|---|---|---|
| Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves) | 37 | 78% | 22% | 0 |
| Real key errors from the September backfill, restored | 28 | 57% | 43% | 0 |
| Keys the backfill disputed, then verified (learner argues the wrong answer) | 35 | 74% | 26% | 0 |
| Explanation from another question | 34 | 91% | 9% | 0 |
| Key moved to a distractor | 70 | 80% | 20% | 0 |
| Sound question, mistaken report | 50 | 76% | 22% | 1 |
| Stem cut off | 35 | 97% | 0% | 1 |
| Option rows disagree with the content key | 14 | 86% | 14% | 0 |
| Image stripped from an image question | 19 | 100% | 0% | 0 |
| Image present, "image missing" report | 16 | 56% | 44% | 0 |
| Linked image is dead in storage | 2 | 100% | 0% | 0 |
Every harmful case, read by hand
The scorer calls a case harmful when the final key differs from the label, a sound question is disabled, or a known defect stays live. Each one was checked against the question itself:
| Case | Course / format | Outcome | Adjudication |
|---|---|---|---|
| Sound question, mistaken report | CFA / fill_blank | fixed | Not harmful. "USD 3.0 million" became 3 with the unit added to the stem; the old key could not be entered on the numeric keypad. |
| Stem cut off | CFA / matrix_grid | no_change | Weak seed. The cut left the matching item answerable from its rows; both solvers answered it with full confidence. |
Real reports: old agent vs new
Labelled real reports only (both independent solvers agree). These are Indian PG MCQs, the old agent's home ground; on every other course and format the old agent did nothing.
| Stored key | Cases | Old agent | New agent |
|---|---|---|---|
| Right (both solvers agree with it) | 18 | left alone 18, changed wrongly 0 | left alone 18, changed wrongly 0 |
| Wrong (both solvers agree on another answer) | 19 | fixed 12, not fixed 7 | fixed 11, disabled 4, to a person 4, missed 0 |
Models on AI Gateway
The same pipeline with each model doing the blind solve and the review, on a stratified 120-case subset (12 per case type). Key changes are confirmed by Claude Opus 5.5 in every run. The three alternatives ran on an earlier build of the checks, so treat the comparison as rough.
| Review + solve model | Scored cases | Correct | To a person | Harmful | Key changes against the label | Cost / review | Median time |
|---|---|---|---|---|---|---|---|
| GPT-6.1-sol (default) | 119 | 81% | 18% | 1 | 1 | $0.039 | 27s |
| GPT-6-sol (fallback) | 119 | 82% | 15% | 3 | 1 | $0.034 | 25s |
| Claude Opus 5.5 | 120 | 78% | 18% | 5 | 2 | $0.128 | 32s |
| Gemini 3.8 Flash | 120 | 69% | 28% | 3 | 1 | $0.035 | 34s |
| GPT-6-luna | 120 | 78% | 18% | 5 | 1 | $0.006 | 21s |
By course and format
| Course / format | Cases | Correct | To a person | Harmful |
|---|---|---|---|---|
| Bar Exam / cloze_dropdown | 9 | 67% | 33% | 0 |
| Bar Exam / integrated_question_set | 2 | 50% | 50% | 0 |
| Bar Exam / matrix_grid | 5 | 80% | 20% | 0 |
| Bar Exam / mcq_single | 14 | 93% | 7% | 0 |
| Bar Exam / ordered_response | 5 | 100% | 0% | 0 |
| Bar Exam / performance_task | 2 | 0% | 100% | 0 |
| Bar Exam / sata | 5 | 100% | 0% | 0 |
| Bar Exam / short_answer | 1 | 100% | 0% | 0 |
| CFA / case_study | 6 | 67% | 33% | 0 |
| CFA / fill_blank | 9 | 67% | 22% | 1 |
| CFA / highlight_hot_spot | 4 | 100% | 0% | 0 |
| CFA / matrix_grid | 5 | 60% | 20% | 1 |
| CFA / mcq_single | 41 | 73% | 27% | 0 |
| CFA / ordered_response | 5 | 80% | 20% | 0 |
| CFA / sata | 11 | 91% | 9% | 0 |
| CPA / cloze_dropdown | 7 | 71% | 29% | 0 |
| CPA / cpa_tbs | 2 | 50% | 50% | 0 |
| CPA / data_entry_grid | 7 | 71% | 29% | 0 |
| CPA / mcq_single | 14 | 93% | 7% | 0 |
| Indian Medical PG / mcq_single | 46 | 80% | 20% | 0 |
| LSAT / argumentative_essay | 1 | 100% | 0% | 0 |
| LSAT / mcq_single | 20 | 100% | 0% | 0 |
| LSAT / reading_comprehension | 2 | 100% | 0% | 0 |
| MCAT / mcq_single | 10 | 60% | 40% | 0 |
| NCLEX-RN / case_study | 6 | 67% | 33% | 0 |
| NCLEX-RN / cloze_dropdown | 10 | 90% | 10% | 0 |
| NCLEX-RN / fill_blank | 9 | 100% | 0% | 0 |
| NCLEX-RN / highlight_hot_spot | 10 | 90% | 10% | 0 |
| NCLEX-RN / image_hot_spot | 5 | 80% | 20% | 0 |
| NCLEX-RN / matrix_grid | 10 | 90% | 10% | 0 |
| NCLEX-RN / mcq_single | 16 | 69% | 31% | 0 |
| NCLEX-RN / ordered_response | 10 | 80% | 20% | 0 |
| NCLEX-RN / sata | 16 | 56% | 44% | 0 |
| UK Medical PG / mcq_single | 5 | 100% | 0% | 0 |
| US Medical PG / mcq_single | 10 | 90% | 10% | 0 |
Every report since 2024
2650 reports (1843 from 480 learners, 807 from the system user) on 2065 questions; 381 questions were reported more than once. 817 learner reports carry a written note, 325 a screenshot. Jev classified every note by what the learner claims (200 name a specific answer; 7 cite a source):
| What the learner claims (Jev) | Notes | Share |
|---|---|---|
| The key is wrong / another answer is right | 295 | 36% |
| Image missing, wrong or unreadable | 171 | 21% |
| A choice is wrong, duplicated or missing | 92 | 11% |
| Stem unclear, incomplete or has a typo | 90 | 11% |
| Note does not say what is wrong | 80 | 10% |
| Explanation wrong or mismatched | 56 | 7% |
| App problem, not the question | 19 | 2% |
| Outdated or off-syllabus | 14 | 2% |
| Old agent action within 3 hours | Reports | Share |
|---|---|---|
| no change | 979 | 37% |
| key changed | 568 | 21% |
| explanation only | 559 | 21% |
| stem edited | 356 | 13% |
| image changed | 144 | 5% |
| disabled | 44 | 2% |
Not in this PR: validating the whole bank
This PR reviews one question at a time, when a learner reports it. A pass over all 113K live question sets is a separate PR: a targeted set of Jev checks per course and question type, run across the bank, with the full review reserved for what Jev flags. Jev (TypeSafe's System One model) told wrong keys from right ones with AUC 0.86 in a first measurement on this benchmark, without reasoning, at about $0.04 per million input tokens.
What else this found
- Fill-in-the-blank grading. The app gives a numeric keypad whenever a key is a number after dropping "$", "," and "%", but the grader kept the "%", so keys like "0.52%" could never be matched. Fixed in the grader in this PR. Keys that embed words ("USD 3.0 million", "163 contracts (sell)") are content defects the review fixes when they are reported.
- Dead images. Some image questions link to files that return 400 from storage; the review records these, redraws the image when the question and explanation say enough to draw it from, and otherwise takes the question out of practice.
- Empty keys. 12 live NCLEX fill-in-the-blank items had an empty answer, so no learner could get them right.