Skip to content
Search
paperJuly 2025Unreviewed

Large language models provide unsafe answers to patient-posed medical questions

R. Draelos, Samina Afreen, Barbara Blasko, Tiffany Brazile, Natasha Chase, Dimple Desai, Jessica Evert, H. Gardner, Lauren Herrmann, A. V. House, Stephanie Kass, Marianne Kavan, Kirshma Khemani, Amanda M. Koire, L. McDonald, Zahraa Rabeeah, A. Shah

npj Digital Medicine

Abstract

Millions of patients are regularly using large language model (LLM) chatbots for medical advice, raising patient safety concerns. This physician-led red-teaming study compares the safety of four publicly available chatbots—Claude by Anthropic, Gemini by Google, GPT-4o by OpenAI, and Llama-3.0/3.1-70B by Meta—on a new dataset, HealthAdvice, using an evaluation framework that enables quantitative and qualitative analysis. In total, 888 chatbot responses are evaluated for 222 patient-posed advice-s

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@article{draelos2025large,
  title = {{Large language models provide unsafe answers to patient-posed medical questions}},
  author = {R. Draelos and Samina Afreen and Barbara Blasko and Tiffany Brazile and Natasha Chase and Dimple Desai and Jessica Evert and H. Gardner and Lauren Herrmann and A. V. House and Stephanie Kass and Marianne Kavan and Kirshma Khemani and Amanda M. Koire and L. McDonald and Zahraa Rabeeah and A. Shah},
  year = {2025},
  month = jul,
  journal = {npj Digital Medicine},
  eprint = {2507.18905},
  archivePrefix = {arXiv},
  doi = {10.1038/s41746-026-02428-5},
  url = {https://www.semanticscholar.org/paper/ebdb2c9347e5ab1b7c27c483aafab9584d28e8a1}
}