Skip to content
Search
paperSeptember 2026Unreviewed

When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems

Wenji Bai, Muhammad Waseem, Zeeshan Rasheed, Jaakko Peltonen, Pekka Abrahamsson

Abstract

LLM-based agents for automated code repair have received significant attention in recent years from both research and software engineering practice perspectives. However, limited attention has been paid to patches that pass syntactic and functional verification but still retain or introduce security vulnerabilities. The aim of this research is to systematically identify and categorize such silent failures in LLM-based agentic code repair. We conducted an empirical study using 1,030 valid executi

Categories

Cite

@misc{bai2026when,
  title = {{When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems}},
  author = {Wenji Bai and Muhammad Waseem and Zeeshan Rasheed and Jaakko Peltonen and Pekka Abrahamsson},
  year = {2026},
  month = sep,
  eprint = {2609.10548},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.10548}
}