Skip to content
Search
paperJune 2026Unreviewed

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems

Leonhard Waibl, Felix Michalak, Hadrien Mariaccia

Abstract

LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost and latency, and rarely compare specialized guardrails against repurposed generalist LLMs. We present BELLS-O (Benchmark for the Evaluation of LLM Supervision Systems, Operational), the first independent operational benchmark of LLM supervision systems. BELLS-O evaluates 28 systems

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{waibl2026bellso,
  title = {{BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems}},
  author = {Leonhard Waibl and Felix Michalak and Hadrien Mariaccia},
  year = {2026},
  month = jun,
  eprint = {2606.20668},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.20668}
}