Skip to content
Search
paperJune 2026Unreviewed

Gate AI: LLM Security Benchmark Evaluation Methodology and Results

Ryle Goehausen, Marcus Sousa

Abstract

Published evaluations of prompt-injection and jailbreak detectors for Large Language Models often suffer from two systematic weaknesses: per-dataset threshold tuning and undisclosed operating points. We describe an evaluation harness that addresses both. The detector under evaluation is scored across 16 public benchmarks (12,111 samples) using 5-fold cross-validation. StratifiedKFold (by row) is the headline pass; a parallel StratifiedGroupKFold pass over a composite key (parent-prompt id plus M

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{goehausen2026gate,
  title = {{Gate AI: LLM Security Benchmark Evaluation Methodology and Results}},
  author = {Ryle Goehausen and Marcus Sousa},
  year = {2026},
  month = jun,
  eprint = {2606.02959},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.02959}
}