Skip to content
Search
paperDecember 2025Unreviewed

Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks

V. Nguyen, M. Husain

arXiv.org

Abstract

Agentic AI introduces security vulnerabilities that traditional LLM safeguards fail to address. Although recent work by Unit 42 at Palo Alto Networks demonstrated that ChatGPT-4o successfully executes attacks as an agent that it refuses in chat mode, there is no comparative analysis in multiple models and frameworks. We conducted the first systematic penetration testing and comparative evaluation of agentic AI systems, testing five prominent models (Claude 3.5 Sonnet, Gemini 2.5 Flash, GPT-4o, G

Categories

Cite

@misc{nguyen2025penetration,
  title = {{Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks}},
  author = {V. Nguyen and M. Husain},
  year = {2025},
  month = dec,
  eprint = {2512.14860},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2512.14860},
  url = {https://www.semanticscholar.org/paper/71a878eaf0ba150a611610e3c5cbe6cebec1528d}
}