Skip to content
Search
paperJune 2026Unreviewed

Evaluation Methods for LLM Safety and Reliability in Clinical and Healthcare Applications

Anand Jha, Kirtiraj Bhatele, Pratyush Mihir

Abstract

Large Language Models (LLMs) are being deployed in clinical and healthcare systems, which requires serious consideration of their safety, reliability. This chapter investigates hallucinations, harms to patients and healthcare employees as well as methods of measuring the performance, deployment obstacles and ethical issues of LLMs in medical applications. The regulatory frameworks, such as FDA and the EU AI act are reviewed to put compliance and governance requirements into perspective.

Categories

Cite

@misc{jha2026evaluation,
  title = {{Evaluation Methods for LLM Safety and Reliability in Clinical and Healthcare Applications}},
  author = {Anand Jha and Kirtiraj Bhatele and Pratyush Mihir},
  year = {2026},
  month = jun,
  doi = {10.4018/979-8-3373-7862-6.ch005},
  url = {https://doi.org/10.4018/979-8-3373-7862-6.ch005}
}