100% synthetic Demonstration & evaluation use only Not clinically validated
THE 60-SECOND EXPLAINER

Clinical coding:
the animated version.

From ambulance bay to OPCS-4—and why British coding AI needs a British test track. Starring one hi-vis pigeon, a surgical rubber duck and a deeply concerned finance department.

01:00 · SOUND-OPTIONAL CINEMA
Read the plain-English transcript

A patient episode begins when someone arrives for care. Doctors and nurses treat the patient and document the encounter in notes, scans and discharge summaries. Across NHS trusts, specialist clinical-coding teams—sometimes more than thirty people in a large trust—read those records and translate confirmed diagnoses and completed procedures into ICD-10 and OPCS-4 codes.

It is skilled, labour-intensive work that takes years to master. The resulting codes support reimbursement, epidemiology and service planning. Automation could help, but patient records are sensitive and difficult for AI companies to obtain. MIMIC-IV gives researchers reusable US hospital data, but it reflects an American institution and coding context rather than NHS workflows.

EPICODE is the MVP of a UK-specific synthetic episode dataset and benchmark for training, comparing and evaluating clinical-coding AI products before governed real-world validation.

The missing UK clinical-coding test layer

Test the code.
Trace the proof.

US datasets made clinical AI evaluation possible—but they encode US care and billing. NHS teams need UK-native episodes mapped to ICD-10 and OPCS-4, without waiting for patient-level data access.

Generate a case
Deterministic Evidence linked Schema validated
Synthetic by construction. Serious by intent.
TRACE VIEW
S
Synthetic adult 026General Surgery · Validated
ASSESSMENT · 09:20

Clinical impression: acute appendicitis.

ICD-10 · 2026 K35.8 Acute appendicitis 2 evidence links
Schema valid100% evidence coverage
Export readyJSON · schema v1
10
Curated scenarios2 clinical tracks
WHY THIS EXISTS

The benchmark gap
is the product.

The UK does not lack health data. It lacks a ready-to-use, public, coherent episode benchmark for testing NHS clinical coding systems against both diagnosis and procedure standards.

01 / USA

Reusable data.
Wrong national context.

MIMIC-IV provides de-identified hospital records, notes and ICD billing fields at research scale. It is an extraordinary resource—and it comes from one Boston health system, not the NHS.

Useful for model research ≠ evidence of NHS coding performance.
02 / UK

Relevant data.
Not a plug-in benchmark.

HES contains nationally coded activity, but patient-level access is governed. Public Artificial HES helps test pipelines, yet does not preserve relationships between fields and cannot support analysis.

No public MIMIC-equivalent for coherent NHS ICD-10 + OPCS-4 episode evaluation.
THE BRIDGE

EPICODE supplies the missing test fixture: deterministic NHS-style episodes, expected ICD-10 and OPCS-4 classifications, and passage-level evidence contracts. Synthetic first; governed real-world validation later.

Inspect the test data ↓
EPISODE FACTORY

From scenario to traceable dossier

Choose a pathway and seed. The same inputs always produce the same validated episode.

01 Choose a clinical track
02 Select a scenario
No API key requiredICD-10 5th Edition 2026OPCS-4.11
THE TRANSATLANTIC DATA SITUATION

The US benchmark is not an NHS benchmark.

America offers reusable clinical datasets at research scale. Britain offers rich national activity data behind governance and public artificial extracts that are useful for plumbing—not coherent coding evaluation. EPICODE is designed for that specific space between them.

UNITED KINGDOM

Treasure behind a drawbridge

Rich, governed

Commercial effect: superb real-world evidence exists, and public artificial HES helps test plumbing. But coherent product evaluation can still inherit access lead-time, governance cost and restricted portability.

£?
THE BUSINESS CASE

Build the test track before renting the motorway.

A curated UK synthetic benchmark lets product, clinical and procurement teams agree acceptance criteria early. Use it for rapid regression tests, evidence demos and vendor comparisons; reserve governed real data for later-stage external validation.

  • FasterStart before data-access approval
  • SaferDevelop without patient exposure
  • CheaperCatch failures before pilots
  • UK-nativeICD-10 + OPCS-4 pathways

Landscape snapshot, 18 July 2026. “Fewer ready-made benchmarks” is our inference from the access models shown—not a claim that UK health data is absent. Synthetic results establish engineering confidence, not clinical validity.

WHY IT MATTERS

A safer starting line for clinical AI evaluation

01

No patient exposure

Purpose-built synthetic records let teams develop evaluation workflows before governed data access.

02

Traceability by design

Every classification points back to explicit evidence, making expected outputs inspectable.

03

Repeatable inputs

Seeded generation turns a polished demo into a reproducible engineering asset.