Skip to content
Search
paperMay 2026Unreviewed

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

Yige Li, Jun Sun, Wei Zhao, Zhe Li, Yutao Wu, Hanxun Huang, Xiang Zheng, Xingjun Ma

Abstract

Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce \textsc{MedHarm}\footnote{Code and data will be released upon acceptance. Due to the sensitive nature of high-risk medical queries, data access will be available to qualified researchers upon request.}, a high-risk medical safety benchmark with 1,100 medically grounded queries across 10 safety-critical categories, inc

Categories

Cite

@misc{li2026whena,
  title = {{When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries}},
  author = {Yige Li and Jun Sun and Wei Zhao and Zhe Li and Yutao Wu and Hanxun Huang and Xiang Zheng and Xingjun Ma},
  year = {2026},
  month = may,
  eprint = {2606.28332},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.28332}
}