arXiv cs.LG
· Papers
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
arXiv:2508.00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against. Here we introduce