September 2026Unreviewed
When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems
Wenji Bai, Muhammad Waseem, Zeeshan Rasheed, Jaakko Peltonen, Pekka Abrahamsson
Abstract
LLM-based agents for automated code repair have received significant attention in recent years from both research and software engineering practice perspectives. However, limited attention has been paid to patches that pass syntactic and functional verification but still retain or introduce security vulnerabilities. The aim of this research is to systematically identify and categorize such silent failures in LLM-based agentic code repair. We conducted an empirical study using 1,030 valid executi
Categories
Cite
@misc{bai2026when,
title = {{When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems}},
author = {Wenji Bai and Muhammad Waseem and Zeeshan Rasheed and Jaakko Peltonen and Pekka Abrahamsson},
year = {2026},
month = sep,
eprint = {2609.10548},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.10548}
}