February 2026Unreviewed
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami, G. Gidel, Stephan Gunnemann
arXiv.org
Abstract
Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate harmfulness in order to benchmark the robustness of safety against adversarial attacks. However, we show that existing validation protocols fail to account for substantial distribution shifts inherent to red-teaming: diverse victim models exhibit distinct generation styles, attacks
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{schwinn2026coin,
title = {{A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness}},
author = {Leo Schwinn and Moritz Ladenburger and Tim Beyer and Mehrnaz Mofakhami and G. Gidel and Stephan Gunnemann},
year = {2026},
month = feb,
eprint = {2603.06594},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2603.06594},
url = {https://www.semanticscholar.org/paper/e8bd6ec12610833e7de437d7b2a5b4fc7c89fd9d}
}