Skip to content
Search
paperJune 2026UnreviewedOpen access

A red teaming framework for large language models: a case study on faithfulness evaluation

Abrar Alotaibi, Raed Mughus, Moataz Ahmed

Software quality journal

Abstract

Large language models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks, yet their deployment in high-stakes applications has raised critical concerns regarding reliability, safety, and response trustworthiness. In this paper, we present a red teaming framework that systematically uncovers vulnerabilities in LLM outputs. Our approach employs a novel multi-role architecture comprising a target, attackers, and jury models. The attackers genera

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@article{alotaibi2026red,
  title = {{A red teaming framework for large language models: a case study on faithfulness evaluation}},
  author = {Abrar Alotaibi and Raed Mughus and Moataz Ahmed},
  year = {2026},
  month = jun,
  journal = {Software quality journal},
  eprint = {2606.25476},
  archivePrefix = {arXiv},
  doi = {10.1007/s11219-026-09779-y},
  url = {https://www.semanticscholar.org/paper/f268fa307779519fcd044adf0b3bc5b8f828611a}
}