June 2026Unreviewed
Gate AI: LLM Security Benchmark Evaluation Methodology and Results
Ryle Goehausen, Marcus Sousa
Abstract
Published evaluations of prompt-injection and jailbreak detectors for Large Language Models often suffer from two systematic weaknesses: per-dataset threshold tuning and undisclosed operating points. We describe an evaluation harness that addresses both. The detector under evaluation is scored across 16 public benchmarks (12,111 samples) using 5-fold cross-validation. StratifiedKFold (by row) is the headline pass; a parallel StratifiedGroupKFold pass over a composite key (parent-prompt id plus M
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{goehausen2026gate,
title = {{Gate AI: LLM Security Benchmark Evaluation Methodology and Results}},
author = {Ryle Goehausen and Marcus Sousa},
year = {2026},
month = jun,
eprint = {2606.02959},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.02959}
}