June 2026UnreviewedOpen access
A red teaming framework for large language models: a case study on faithfulness evaluation
Abrar Alotaibi, Raed Mughus, Moataz Ahmed
Software quality journal
Abstract
Large language models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks, yet their deployment in high-stakes applications has raised critical concerns regarding reliability, safety, and response trustworthiness. In this paper, we present a red teaming framework that systematically uncovers vulnerabilities in LLM outputs. Our approach employs a novel multi-role architecture comprising a target, attackers, and jury models. The attackers genera
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@article{alotaibi2026red,
title = {{A red teaming framework for large language models: a case study on faithfulness evaluation}},
author = {Abrar Alotaibi and Raed Mughus and Moataz Ahmed},
year = {2026},
month = jun,
journal = {Software quality journal},
eprint = {2606.25476},
archivePrefix = {arXiv},
doi = {10.1007/s11219-026-09779-y},
url = {https://www.semanticscholar.org/paper/f268fa307779519fcd044adf0b3bc5b8f828611a}
}