Open clinical AI evaluation

Evidence first.
Confidence measured.

An Ollama-first workbench for testing structured extraction, citation support, abstention, and calibration on clinical information tasks.

Run it locally Inspect the real model report Request an evaluation sprint
Starter v1

A small benchmark you can audit case by case.

12

Wholly synthetic cases. No patient records or restricted clinical text.

6 pairs

Each complete passage has a near-neighbor with missing or conflicting evidence.

3 domains

Melanoma pathology, transplant evidence, and randomized-trial abstraction.

0 keys

The default run stays local through Ollama and needs no paid model API.

Audit trail

Every score traces back to a raw response.

01Frozen case and prompt
02Local model inference
03Raw JSONL trace
04Deterministic scoring
05Portable error report

What the score does not mean

This repository validates software behavior on a tiny synthetic set. It does not establish medical knowledge, clinical safety, diagnostic performance, or readiness for patient care. A serious research release needs independent domain annotation, adjudication, a sealed test set, and a pre-specified analysis.