Wholly synthetic cases. No patient records or restricted clinical text.
Open clinical AI evaluation
Evidence first.
Confidence measured.
An Ollama-first workbench for testing structured extraction, citation support, abstention, and calibration on clinical information tasks.
Run it locally Inspect the real model report Request an evaluation sprintStarter v1
A small benchmark you can audit case by case.
Each complete passage has a near-neighbor with missing or conflicting evidence.
Melanoma pathology, transplant evidence, and randomized-trial abstraction.
The default run stays local through Ollama and needs no paid model API.
Audit trail
Every score traces back to a raw response.
01Frozen case and prompt
02Local model inference
03Raw JSONL trace
04Deterministic scoring
05Portable error report
What the score does not mean
This repository validates software behavior on a tiny synthetic set. It does not establish medical knowledge, clinical safety, diagnostic performance, or readiness for patient care. A serious research release needs independent domain annotation, adjudication, a sealed test set, and a pre-specified analysis.