Real-World Evaluation of Large Language Models in Healthcare (RWE-LLM): A New Realm of AI Safety & Validation
MD MHA¹ Meenesh Bhimani, BS¹ Alex Miller, P. M. Jonathan D. Agnew, Markel Sanz Ausin, BA¹ Mariska Raglow-Defranco, MD Mba Harpreet Mangat, Bsn RN Ccm Michelle Voisard, RN Bsn Ccm Maggie Taylor, BS Sebastian Bierman-Lytle, BS BA Vishal Parikh, Juliana Ghukasyan, BS Rae Lasko, Saad Godil, Meng, MD Mph Ashish Atreja, PhD¹ Subhabrata Mukherjee
medRxiv
Abstract
Background: The deployment of artificial intelligence (AI) in healthcare necessitates robust safety validation frameworks, particularly for systems directly interacting with patients. While theoretical frameworks exist, there remains a critical gap between abstract principles and practical implementation. Traditional LLM benchmarking approaches provide very limited output coverage and are insufficient for healthcare applications requiring high safety standards. Objective: To develop and evaluate
Categories
Cite
@article{bhimani2025realworld,
title = {{Real-World Evaluation of Large Language Models in Healthcare (RWE-LLM): A New Realm of AI Safety \& Validation}},
author = {MD MHA¹ Meenesh Bhimani and BS¹ Alex Miller and P. M. Jonathan D. Agnew and Markel Sanz Ausin and BA¹ Mariska Raglow-Defranco and MD Mba Harpreet Mangat and Bsn RN Ccm Michelle Voisard and RN Bsn Ccm Maggie Taylor and BS Sebastian Bierman-Lytle and BS BA Vishal Parikh and Juliana Ghukasyan and BS Rae Lasko and Saad Godil and Meng and MD Mph Ashish Atreja and PhD¹ Subhabrata Mukherjee},
year = {2025},
month = mar,
journal = {medRxiv},
doi = {10.1101/2025.03.17.25324157},
url = {https://www.semanticscholar.org/paper/1fe644e3b971a52cc57b1b5ddc4cba5408cf1a8d}
}