May 2026Unreviewed
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
Vivek Dahiya, Sunny Nehra, Vipul Dholariya, Bhavik Shangari, Chandra Khatri
Abstract
We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production-style applications with 118 ground-truth vulnerabilities across 20+ CWE families, which we will open-source). We test six frontier models (GPT-5.4, Codex~5.3, Claude Opus~4.6, Sonnet~4.6, Gemini~3.1~Pro and Gemini~3~Flash) and two domain-specialized models acr
Categories
Cite
@misc{dahiya2026are,
title = {{Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks}},
author = {Vivek Dahiya and Sunny Nehra and Vipul Dholariya and Bhavik Shangari and Chandra Khatri},
year = {2026},
month = may,
eprint = {2605.23243},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.23243}
}