June 2026Unreviewed
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
Leonhard Waibl, Felix Michalak, Hadrien Mariaccia
Abstract
LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost and latency, and rarely compare specialized guardrails against repurposed generalist LLMs. We present BELLS-O (Benchmark for the Evaluation of LLM Supervision Systems, Operational), the first independent operational benchmark of LLM supervision systems. BELLS-O evaluates 28 systems
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{waibl2026bellso,
title = {{BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems}},
author = {Leonhard Waibl and Felix Michalak and Hadrien Mariaccia},
year = {2026},
month = jun,
eprint = {2606.20668},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.20668}
}