December 2025Unreviewed
Toward Trustworthy Chatbots: A Protocol for Red Teaming for Health Related Conversations
Syed-Amad Hussain, Daniel I. Jackson, Ashley Lewis, Eric Fosler-Lussier, Emre Sezgin
medRxiv
Abstract
Introduction: Health-related chatbots are increasingly used to mediate conversations that carry clinical significance and emotional weight. Retrieval-augmented generation (RAG) can reduce factual errors ('hallucinations'), but the risks remain, with additional challenges coming from chatbots acting against behavioral safety and scope rules. Red teaming, an adversarial testing process that deliberately probes systems for failures before deployment, offers a way to surface potential risks. We desc
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@article{hussain2025trustworthy,
title = {{Toward Trustworthy Chatbots: A Protocol for Red Teaming for Health Related Conversations}},
author = {Syed-Amad Hussain and Daniel I. Jackson and Ashley Lewis and Eric Fosler-Lussier and Emre Sezgin},
year = {2025},
month = dec,
journal = {medRxiv},
doi = {10.64898/2025.12.15.25342297},
url = {https://www.semanticscholar.org/paper/930205acf84599560a761027e57b15506d11fe7b}
}