December 2025Unreviewed
Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
V. Nguyen, M. Husain
arXiv.org
Abstract
Agentic AI introduces security vulnerabilities that traditional LLM safeguards fail to address. Although recent work by Unit 42 at Palo Alto Networks demonstrated that ChatGPT-4o successfully executes attacks as an agent that it refuses in chat mode, there is no comparative analysis in multiple models and frameworks. We conducted the first systematic penetration testing and comparative evaluation of agentic AI systems, testing five prominent models (Claude 3.5 Sonnet, Gemini 2.5 Flash, GPT-4o, G
Categories
Cite
@misc{nguyen2025penetration,
title = {{Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks}},
author = {V. Nguyen and M. Husain},
year = {2025},
month = dec,
eprint = {2512.14860},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2512.14860},
url = {https://www.semanticscholar.org/paper/71a878eaf0ba150a611610e3c5cbe6cebec1528d}
}