July 2026Unreviewed
Towards an Automated Test of LLM Security Knowledge
Shufan Chai, Liangliang Sun, Jessica Staddon
Abstract
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security ``knowledge'' may be insufficient. Popular strategies for identifying LLM security knowledge gaps include building corpora of challenge questions or task benchmarks, strategies that require substantial manual work and secur
Categories
Cite
@misc{chai2026automated,
title = {{Towards an Automated Test of LLM Security Knowledge}},
author = {Shufan Chai and Liangliang Sun and Jessica Staddon},
year = {2026},
month = jul,
eprint = {2607.18496},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.18496}
}