Skip to content
Search
paperMarch 2026Unreviewed

Benchmarking the Effectiveness of AI-Driven Red Teaming Across Safety-Aligned Language Models

Joyce Malicha, Kamrul Hasan

2026 IEEE 2nd International Conference on Secure IoT, Assured and Trusted Computing (SATC)

Abstract

Large language models (LLMs) have advanced rapidly, yet even safety-aligned models remain vulnerable to adversarial prompts that bypass safeguards and induce harmful outputs. Conventional red teaming methods, including static testing and gradient-based attacks, are limited by either insufficient adaptability or restrictive access assumptions. In this work, we benchmark an autonomous AI-driven evolutionary red teaming system that iteratively generates, evaluates, and refines prompts to uncover sa

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@inproceedings{malicha2026benchmarking,
  title = {{Benchmarking the Effectiveness of AI-Driven Red Teaming Across Safety-Aligned Language Models}},
  author = {Joyce Malicha and Kamrul Hasan},
  year = {2026},
  month = mar,
  booktitle = {2026 IEEE 2nd International Conference on Secure IoT, Assured and Trusted Computing (SATC)},
  doi = {10.1109/SATC69565.2026.11542221},
  url = {https://www.semanticscholar.org/paper/570f0f9f9c7e3b686aa1dc4e97d74ceabe30d149}
}