arXiv cs.CL
· Papers
MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection
arXiv:2607.15166v1 Announce Type: cross Abstract: Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and saf