Search papers, labs, and topics across Lattice.
MedFailBench introduces a novel benchmark designed to assess the safety boundaries of medical AI systems by identifying specific types of failures rather than merely evaluating correct answers. This clinician-built resource includes a comprehensive failure atlas that categorizes errors by severity and type, allowing for a more nuanced understanding of potential risks in AI applications within healthcare. The benchmark currently features 44 synthetic cases with detailed annotations, enabling researchers to systematically analyze and improve AI safety protocols.
Medical AI systems often fail in specific ways, and MedFailBench reveals the critical safety boundaries that need attention.
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains 44 clinician-reviewed synthetic cases with severity annotations, a live HuggingFace leaderboard preview, a safety gate taxonomy, a clinical severity rubric, and an automated pipeline for archiving model-response screening runs. No patient data, clinical validation claims, or model rankings are included. MedFailBench is released under Apache-2.0 and CC-BY-4.0 and carries the Zenodo DOI 10.5281/zenodo.21205535.