May 2026Unreviewed
When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
Yige Li, Jun Sun, Wei Zhao, Zhe Li, Yutao Wu, Hanxun Huang, Xiang Zheng, Xingjun Ma
Abstract
Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce \textsc{MedHarm}\footnote{Code and data will be released upon acceptance. Due to the sensitive nature of high-risk medical queries, data access will be available to qualified researchers upon request.}, a high-risk medical safety benchmark with 1,100 medically grounded queries across 10 safety-critical categories, inc
Categories
Cite
@misc{li2026whena,
title = {{When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries}},
author = {Yige Li and Jun Sun and Wei Zhao and Zhe Li and Yutao Wu and Hanxun Huang and Xiang Zheng and Xingjun Ma},
year = {2026},
month = may,
eprint = {2606.28332},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.28332}
}