September 2026Unreviewed
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, P. Harshangi
Abstract
Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step agent vulnerabilities. We present a systematic black-box framework for risk-aware agent evaluation requiring only basic system descriptions. Our approach introduces: (1) a seven-domain taxonomy mapping observable behaviors to ris
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{kumar2026blackbox,
title = {{Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery}},
author = {Divyanshu Kumar and Nitin Aravind Birur and Tanay Baswa and Sahil Agarwal and P. Harshangi},
year = {2026},
month = sep,
eprint = {2609.09647},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/cf0bc84d59f2f81ec5aa3534b5933d67a80205ec}
}