August 2026Unreviewed
Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code
Animesh Shaw
Abstract
Large language models are increasingly used to author Infrastructure-as-Code (IaC), where a single insecure default can be deployed directly into production. Prior evaluations report raw vulnerability counts for model-generated IaC, but without a human baseline they cannot determine whether models are actually worse than engineers. We introduce GenIaC-SecBench, a benchmark of 100 deployment scenarios stratified by architectural complexity, evaluated across 12 model configurations from four vendo
Categories
Cite
@misc{shaw2026compared,
title = {{Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code}},
author = {Animesh Shaw},
year = {2026},
month = aug,
eprint = {2608.28021},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.28021}
}