Skip to content
Search
paperJuly 2026Unreviewed

Towards an Automated Test of LLM Security Knowledge

Shufan Chai, Liangliang Sun, Jessica Staddon

Abstract

Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security ``knowledge'' may be insufficient. Popular strategies for identifying LLM security knowledge gaps include building corpora of challenge questions or task benchmarks, strategies that require substantial manual work and secur

Categories

Cite

@misc{chai2026automated,
  title = {{Towards an Automated Test of LLM Security Knowledge}},
  author = {Shufan Chai and Liangliang Sun and Jessica Staddon},
  year = {2026},
  month = jul,
  eprint = {2607.18496},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.18496}
}