March 2026Unreviewed
Benchmarking the Effectiveness of AI-Driven Red Teaming Across Safety-Aligned Language Models
Joyce Malicha, Kamrul Hasan
2026 IEEE 2nd International Conference on Secure IoT, Assured and Trusted Computing (SATC)
Abstract
Large language models (LLMs) have advanced rapidly, yet even safety-aligned models remain vulnerable to adversarial prompts that bypass safeguards and induce harmful outputs. Conventional red teaming methods, including static testing and gradient-based attacks, are limited by either insufficient adaptability or restrictive access assumptions. In this work, we benchmark an autonomous AI-driven evolutionary red teaming system that iteratively generates, evaluates, and refines prompts to uncover sa
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@inproceedings{malicha2026benchmarking,
title = {{Benchmarking the Effectiveness of AI-Driven Red Teaming Across Safety-Aligned Language Models}},
author = {Joyce Malicha and Kamrul Hasan},
year = {2026},
month = mar,
booktitle = {2026 IEEE 2nd International Conference on Secure IoT, Assured and Trusted Computing (SATC)},
doi = {10.1109/SATC69565.2026.11542221},
url = {https://www.semanticscholar.org/paper/570f0f9f9c7e3b686aa1dc4e97d74ceabe30d149}
}