Paired challenge cases
Ten standard episodes plus matched traps where diagnoses are unconfirmed and procedures are not performed.
A narrow test of the gap US datasets cannot fill: 20 NHS-style episodes, fourteen UK ICD-10/OPCS-4 codes, and passage-level proof for every answer.
| Rank | Model | Episode resolved | Code F1 | Evidence F1 | Trap pass | Hallucination | Avg latency |
|---|
Set OPENAI_API_KEY, then run .venv/bin/python -m synthetic_episode_studio.benchmark run. Results appear here automatically.
This is an engineering smoke test, not evidence of clinical safety or broad model capability.
Ten standard episodes plus matched traps where diagnoses are unconfirmed and procedures are not performed.
Models see clinical passages and 14 permitted code identifiers—not code descriptions, gold assignments or rationales.
A case resolves only when codes and every supporting passage ID match the gold contract.
Fixed cases, versioned prompt hash, stored raw predictions, token counts and latency.
More benchmark theatre per square metre than strictly necessary.