June 2026Unreviewed
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing
Alexandre Cristovão Maiorano
Abstract
Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat. We measure attribution. We add four OWASP-LLM-Top-10-aware agents to a 21-agent baseline scanner and target a lattice of four synthetic LLM endpoints: $L_0$ (no defenses), $L_1$ (refusa
Categories
Cite
@misc{maiorano2026which,
title = {{Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing}},
author = {Alexandre Cristovão Maiorano},
year = {2026},
month = jun,
eprint = {2606.02822},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.02822}
}