20 synthetic episodes14-code closed worldZero patient data
EEPICODE-Bench Mini
UK CLINICAL CODING EVALv1.1.0

Clinical coding.
Measured.

A narrow test of the gap US datasets cannot fill: 20 NHS-style episodes, fourteen UK ICD-10/OPCS-4 codes, and passage-level proof for every answer.

Read the results v1.1.0 · 20 cases
EPICODE / v1.1.0COMPLETE
STRICT RESOLUTION10%
CODE F1100%
EVIDENCE F176.9%
02 MODELS20 EPISODESNO PATIENT DATA
THE MAIN EVENT

EPICODE-Bench Mini leaderboard

Primary metric Episode Resolved requires the exact code set and exact supporting passage set.
RankModelEpisode resolvedCode F1Evidence F1Trap passHallucinationAvg latency
FrontierCost-sensitive↳ Same prompt · Same ten cases · Model-specific reasoning
RULES OF ENGAGEMENT

Small enough to understand.
Strict enough to fail.

This is an engineering smoke test, not evidence of clinical safety or broad model capability.

01

Paired challenge cases

Ten standard episodes plus matched traps where diagnoses are unconfirmed and procedures are not performed.

02

Descriptions stay hidden

Models see clinical passages and 14 permitted code identifiers—not code descriptions, gold assignments or rationales.

03

Exact means exact

A case resolves only when codes and every supporting passage ID match the gold contract.

04

Reproducible by design

Fixed cases, versioned prompt hash, stored raw predictions, token counts and latency.

20episodes
2clinical tracks
14candidate codes
40model calls

More benchmark theatre per square metre than strictly necessary.