Skip to content
Search
paperAugust 2026Unreviewed

Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code

Animesh Shaw

Abstract

Large language models are increasingly used to author Infrastructure-as-Code (IaC), where a single insecure default can be deployed directly into production. Prior evaluations report raw vulnerability counts for model-generated IaC, but without a human baseline they cannot determine whether models are actually worse than engineers. We introduce GenIaC-SecBench, a benchmark of 100 deployment scenarios stratified by architectural complexity, evaluated across 12 model configurations from four vendo

Categories

Cite

@misc{shaw2026compared,
  title = {{Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code}},
  author = {Animesh Shaw},
  year = {2026},
  month = aug,
  eprint = {2608.28021},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.28021}
}