Skip to content
Search
paperJune 2026Unreviewed

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

Alexandre Cristovão Maiorano

Abstract

Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat. We measure attribution. We add four OWASP-LLM-Top-10-aware agents to a 21-agent baseline scanner and target a lattice of four synthetic LLM endpoints: $L_0$ (no defenses), $L_1$ (refusa

Categories

Cite

@misc{maiorano2026which,
  title = {{Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing}},
  author = {Alexandre Cristovão Maiorano},
  year = {2026},
  month = jun,
  eprint = {2606.02822},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.02822}
}