February 2026Unreviewed
Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming
I. Steenstra, Paola Pedrelli, Weiyan Shi, S. Marsella, Timothy W. Bickmore
arXiv.org
Abstract
Large Language Models (LLMs) are increasingly utilized for mental health support; however, current safety benchmarks often fail to detect the complex, longitudinal risks inherent in therapeutic dialogue. We introduce an evaluation framework that pairs AI psychotherapists with simulated patient agents equipped with dynamic cognitive-affective models and assesses therapy session simulations against a comprehensive quality of care and risk ontology. We apply this framework to a high-impact test cas
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{steenstra2026assessing,
title = {{Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming}},
author = {I. Steenstra and Paola Pedrelli and Weiyan Shi and S. Marsella and Timothy W. Bickmore},
year = {2026},
month = feb,
eprint = {2602.19948},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2602.19948},
url = {https://www.semanticscholar.org/paper/98a12d04ba9a6d07bab2a443b7994dd805eeea35}
}