Question review agent
ONC2-7195 ยท Replaces the report fixer (fix-question-task-v2) with one pipeline for every course and format. All numbers from prod data, read-only, on 2026-09-28.
What was wrong
- New formats never got reviewed. The fixer threw for every format except
mcq_singleand skipped every non-medical course, so reports on CFA, CPA, LSAT, Bar and NCLEX's new formats failed or were ignored. A report also sets the question toinvalid, and several practice queries only servevalidatedquestions, so a failed or skipped review took the question out of those pools for good. - It edited the wrong fields. New formats render and grade from
content; the old agent only editedquestion_text, the option rows and the explanation, so a key it "fixed" on a new-format question kept grading against the old key. - Its key changes were unchecked. It changed the key after 568 of 2650 reports (21%). 67 questions had their key changed after two separate reports, and in 30 of them the second fix undid the first.
- Old machinery. A Claude Agent SDK subprocess with six database-writing tools and a medical-only prompt on Sonnet 4.6; no benchmark.
How it works now
- Load the whole question. The reported row plus its set (case study, TBS, question set, reading set), hidden parts dropped, images checked to be reachable.
- Code checks the structure. Per-format checks, and the stored key is submitted to the real grader and must score 100%. Code findings are facts the review must resolve.
- Blind solve. GPT-6-sol answers every keyed part with the key, explanation and report withheld. (The approach the September explanation backfill used for key disputes.)
- Review agent. An AI SDK
ToolLoopAgenton GPT-6-sol via AI Gateway, with a calculator and Perplexity search, judges the report against the key, the blind answer and the course's exam conventions, and returns a typed verdict: no change, fix, disable or escalate. - Deterministic, checked writes. Fixes are exact text replacements or a new key in the format's own terms, written into every copy (content, option rows, flattened text, explanation copies). A fix goes live only if: a key changes only when the blind solve of the original question (at 0.6+ confidence) or a structural check disagreed with it; every changed key is reached by two blind solvers from different model families (GPT-6-sol and Claude Opus 5.5) at 0.8+ confidence; every part whose text changed is re-solved from the new text and still reaches its key; the question lints clean; and the rebuilt explanation passes the house-format checks. Disabling needs a structural finding or a confident blind-solve disagreement on the item the reviewer blames. Otherwise a person decides.
- Record and notify. One transaction that refuses a question edited since it was read, never moves a disabled question back into practice, and keeps the previous fields and status in
metadata.question_reviewso a run can be reverted; then audit rows, a Slack summary and the learner email.
Vercel's eve was considered and not used: it is a separate agent server with its own workflow store, sandbox and auth, which would duplicate Trigger.dev for a job that is one bounded review per report.
Benchmark: 363 frozen cases
Built read-only from prod (scripts/question-review/build-benchmark.ts, data in scripts/question-review/data/benchmark.jsonl.gz). Scored per case as correct, to a person (held or escalated when a fix or no change was expected) or harmful (a wrong key applied, a sound question disabled, or a known defect left live).
| Case type | Cases | Correct | To a person | Harmful |
|---|---|---|---|---|
| Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves) | 37 | 78% | 11% | 4 |
| Real key errors from the September backfill, restored | 28 | 46% | 54% | 0 |
| Keys the backfill disputed, then verified (learner argues the wrong answer) | 35 | 71% | 23% | 2 |
| Sound question, mistaken report | 50 | 80% | 18% | 1 |
| Key moved to a distractor | 70 | 81% | 19% | 0 |
| Explanation from another question | 34 | 91% | 9% | 0 |
| Stem cut off | 35 | 97% | 0% | 1 |
| Option rows disagree with the content key | 14 | 86% | 14% | 0 |
| Image stripped from an image question | 19 | 100% | 0% | 0 |
| Image present, "image missing" report | 16 | 62% | 31% | 1 |
| Linked image is dead in storage | 2 | 100% | 0% | 0 |
Every harmful case, read by hand
The scorer calls a case harmful when the final key differs from the label, a sound question is disabled, or a known defect stays live. Each one was checked against the question itself:
| Case | Course / format | Outcome | Adjudication |
|---|---|---|---|
| Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves) | Indian Medical PG / mcq_single | fixed | Likely a real miss. Indian PG texts describe flat vegetations in valve pockets as Libman-Sacks, as both labelling models answered; the review's own solver agreed with the stored NBTE key, so the review kept it and rebuilt the explanation. Needs a person. |
| Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves) | Indian Medical PG / mcq_single | fixed | Fine. The labelled answer was a second correct choice; the review replaced it with a true restrictive operation and tightened the stem, and both solvers reach the key. |
| Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves) | Indian Medical PG / mcq_single | fixed | Judgment call. Correct now, but the review narrowed the stem to keep the disputed key rather than changing the key. |
| Real learner reports (key labelled by GPT-6-astra + Claude Opus 5.5 blind solves) | Indian Medical PG / mcq_single | fixed | Judgment call. Correct now, but the review narrowed the stem to keep the disputed key rather than changing the key. |
| Keys the backfill disputed, then verified (learner argues the wrong answer) | CFA / sata | fixed | Not harmful. The same three statements stay keyed; one was reworded to be accurate, which the scorer counts as a key change. |
| Keys the backfill disputed, then verified (learner argues the wrong answer) | CFA / mcq_single | fixed | Not harmful. It fixed a distractor that was a second correct answer, and turned a sibling part's key "USD 8,920,000" into 8920000, the same value in a form learners can enter. |
| Stem cut off | CFA / matrix_grid | no_change | Weak seed. The cut left the matching item answerable from its rows; both solvers answered it with full confidence. |
| Sound question, mistaken report | CFA / fill_blank | fixed | Not harmful. "USD 3.0 million" became 3 with the unit added to the stem; the old key could not be entered on the numeric keypad. |
| Image present, "image missing" report | CFA / case_study | disabled | Real defect the label missed. The figure's one- and two-year spot rates give a forward rate near 4.00%, but the key is 4.21%, so the case was rightly taken out of practice. |
Real reports: old agent vs new
Labelled real reports only (both independent solvers agree). These are Indian PG MCQs, the old agent's home ground; on every other course and format the old agent did nothing.
| Stored key | Cases | Old agent | New agent |
|---|---|---|---|
| Right (both solvers agree with it) | 18 | left alone 18, changed wrongly 0 | left alone 18, changed wrongly 0 |
| Wrong (both solvers agree on another answer) | 19 | fixed 12, not fixed 7 | fixed 11, disabled 2, to a person 2, missed 4 |
Models on AI Gateway
The same pipeline with each model doing the blind solve and the review, on a stratified 120-case subset (12 per case type). Key changes are confirmed by Claude Opus 5.5 in every run. The three alternatives ran on an earlier build of the checks, so treat the comparison as rough.
| Review + solve model | Scored cases | Correct | To a person | Harmful | Key changes against the label | Cost / review | Median time |
|---|---|---|---|---|---|---|---|
| GPT-6-sol (default) | 119 | 82% | 15% | 3 | 1 | $0.034 | 25s |
| Claude Opus 5.5 | 120 | 78% | 18% | 5 | 2 | $0.128 | 32s |
| Gemini 3.8 Flash | 120 | 69% | 28% | 3 | 1 | $0.035 | 34s |
| GPT-6-luna | 120 | 78% | 18% | 5 | 1 | $0.006 | 21s |
By course and format
| Course / format | Cases | Correct | To a person | Harmful |
|---|---|---|---|---|
| Bar Exam / cloze_dropdown | 9 | 78% | 22% | 0 |
| Bar Exam / integrated_question_set | 2 | 100% | 0% | 0 |
| Bar Exam / matrix_grid | 5 | 100% | 0% | 0 |
| Bar Exam / mcq_single | 14 | 93% | 7% | 0 |
| Bar Exam / ordered_response | 5 | 100% | 0% | 0 |
| Bar Exam / performance_task | 2 | 50% | 50% | 0 |
| Bar Exam / sata | 5 | 80% | 20% | 0 |
| Bar Exam / short_answer | 1 | 100% | 0% | 0 |
| CFA / case_study | 6 | 67% | 17% | 1 |
| CFA / fill_blank | 9 | 67% | 22% | 1 |
| CFA / highlight_hot_spot | 4 | 100% | 0% | 0 |
| CFA / matrix_grid | 5 | 60% | 20% | 1 |
| CFA / mcq_single | 41 | 68% | 29% | 1 |
| CFA / ordered_response | 5 | 80% | 20% | 0 |
| CFA / sata | 11 | 82% | 9% | 1 |
| CPA / cloze_dropdown | 7 | 86% | 14% | 0 |
| CPA / cpa_tbs | 2 | 50% | 50% | 0 |
| CPA / data_entry_grid | 7 | 57% | 43% | 0 |
| CPA / mcq_single | 14 | 93% | 7% | 0 |
| Indian Medical PG / mcq_single | 46 | 80% | 11% | 4 |
| LSAT / argumentative_essay | 1 | 100% | 0% | 0 |
| LSAT / mcq_single | 20 | 95% | 5% | 0 |
| LSAT / reading_comprehension | 2 | 100% | 0% | 0 |
| MCAT / mcq_single | 10 | 70% | 30% | 0 |
| NCLEX-RN / case_study | 6 | 67% | 33% | 0 |
| NCLEX-RN / cloze_dropdown | 10 | 90% | 10% | 0 |
| NCLEX-RN / fill_blank | 9 | 89% | 11% | 0 |
| NCLEX-RN / highlight_hot_spot | 10 | 90% | 10% | 0 |
| NCLEX-RN / image_hot_spot | 5 | 100% | 0% | 0 |
| NCLEX-RN / matrix_grid | 10 | 90% | 10% | 0 |
| NCLEX-RN / mcq_single | 16 | 75% | 25% | 0 |
| NCLEX-RN / ordered_response | 10 | 80% | 20% | 0 |
| NCLEX-RN / sata | 16 | 62% | 38% | 0 |
| UK Medical PG / mcq_single | 5 | 80% | 20% | 0 |
| US Medical PG / mcq_single | 10 | 80% | 20% | 0 |
Screening every question
A full review costs about $0.035, so the sweep screens first and reviews only what is flagged. Measured on the benchmark's key errors (131) and sound keys (187):
| Screen | Key errors caught | Sound keys flagged |
|---|---|---|
| gpt-6-luna blind solve | 85.5% | 6.4% |
| Jev >= 0.3 | 75.6% | 7% |
| gpt-6-luna blind solve or Jev >= 0.3 | 87% | 11.2% |
| gpt-6-luna blind solve or Jev >= 0.5 | 85.5% | 7.5% |
| gpt-6-sol blind solve | 88.5% | 3.7% |
GPT-6-luna costs about $0.00023 per question. Its misses are almost all key-copy mismatches, which the code checks catch, so lint plus the luna solve catch about 96% of key errors. Jev (TypeSafe's System One model) scores AUC 0.86 at telling wrong keys from right ones without reasoning; it misses calculation items, so it adds the non-key defect signal rather than being the key screen. Learners' answer data adds a third signal: 181 questions with 30+ answers where one distractor is picked more than twice as often as the key.
Every report since 2024
2650 reports (1843 from 480 learners, 807 from the system user) on 2065 questions; 381 questions were reported more than once. 817 learner reports carry a written note, 325 a screenshot. Jev classified every note by what the learner claims (200 name a specific answer; 7 cite a source):
| What the learner claims (Jev) | Notes | Share |
|---|---|---|
| The key is wrong / another answer is right | 295 | 36% |
| Image missing, wrong or unreadable | 171 | 21% |
| A choice is wrong, duplicated or missing | 92 | 11% |
| Stem unclear, incomplete or has a typo | 90 | 11% |
| Note does not say what is wrong | 80 | 10% |
| Explanation wrong or mismatched | 56 | 7% |
| App problem, not the question | 19 | 2% |
| Outdated or off-syllabus | 14 | 2% |
| Old agent action within 3 hours | Reports | Share |
|---|---|---|
| no change | 979 | 37% |
| key changed | 568 | 21% |
| explanation only | 559 | 21% |
| stem edited | 356 | 13% |
| image changed | 144 | 5% |
| disabled | 44 | 2% |
Sweep of live questions (partial)
The sweep was stopped after CFA and most of CPA to save gateway credit; nothing from it has been written to prod. Its fixes were proposed while the checks were still being tightened, so the full sweep will review them again before anything is applied.
| Course | Live sets | Screened | Passed screen | Reviewed, sound | Fixed | Disabled | To a person |
|---|---|---|---|---|---|---|---|
| CFA | 1808 | 1789 | 1587 | 73 | 93 | 0 | 36 |
| CPA | 1680 | 1465 | 1402 | 39 | 16 | 4 | 4 |
Proposed fixes: 90 key corrections (mostly CFA numeric keys stored with units, such as "USD 360,000" or "4.0x", that learners could not enter), 108 text corrections and 198 rebuilt explanations. Screen cost $1.0, review cost $12.19 for 3254 sets.
What else this found
- Fill-in-the-blank grading. The app gives a numeric keypad whenever a key is a number after dropping "$", "," and "%", but the grader kept the "%", so keys like "0.52%" could never be matched. Fixed in the grader in this PR. Keys that embed words ("USD 3.0 million", "163 contracts (sell)") are content defects the sweep fixes.
- Dead images. Some image questions link to files that return 400 from storage; the review now records these and takes the question out of practice when it depends on the image.
- Empty keys. 12 live NCLEX fill-in-the-blank items had an empty answer, so no learner could get them right.