Skip to content
Search
paperMay 2026Unreviewed

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

Vivek Dahiya, Sunny Nehra, Vipul Dholariya, Bhavik Shangari, Chandra Khatri

Abstract

We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production-style applications with 118 ground-truth vulnerabilities across 20+ CWE families, which we will open-source). We test six frontier models (GPT-5.4, Codex~5.3, Claude Opus~4.6, Sonnet~4.6, Gemini~3.1~Pro and Gemini~3~Flash) and two domain-specialized models acr

Categories

Cite

@misc{dahiya2026are,
  title = {{Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks}},
  author = {Vivek Dahiya and Sunny Nehra and Vipul Dholariya and Bhavik Shangari and Chandra Khatri},
  year = {2026},
  month = may,
  eprint = {2605.23243},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.23243}
}