Skip to content
Search
paperSeptember 2026Unreviewed

VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities

Jiahao Shi, Edward Tsien, Yifeng Di, Hongjiao Zhang, Yuan Tang, Ronit Dey, Ilona Shishov, Gal Netanel, Zvi Grinberg, Vladimir Belousov, Bat-Zion Rotman, Ilan Pinto, Tianyi Zhang

Abstract

The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a vulnerable dependency is actually exploitable. Security analysts typically spend substantial time assessing vulnerability exploitability case by case. Recent LLM agents have emerged as promising candidates for this task gi

Categories

Framework mappings

MITRE ATLAS
  • AML.T0010AI Supply Chain Compromise

Suggested from the entry's categories.

Cite

@misc{shi2026vexbench,
  title = {{VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities}},
  author = {Jiahao Shi and Edward Tsien and Yifeng Di and Hongjiao Zhang and Yuan Tang and Ronit Dey and Ilona Shishov and Gal Netanel and Zvi Grinberg and Vladimir Belousov and Bat-Zion Rotman and Ilan Pinto and Tianyi Zhang},
  year = {2026},
  month = sep,
  eprint = {2609.08040},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.08040}
}