Skip to content
Search
paperJune 2026Unreviewed

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs

Vincent Limbach, Jonas Dornbusch, David Lüdke, Stephan Günnemann, Leo Schwinn

Abstract

Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense comparison unreliable. Historically, standardized attacks such as AutoAttack have largely resolved this for image classifiers, providing a reliable evaluation baseline for systematic comparison across defenses. However, no equivalent exists for LLM jailbreak evaluation yet, where designing such an attack is considerably m

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{limbach2026blackbox,
  title = {{Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs}},
  author = {Vincent Limbach and Jonas Dornbusch and David Lüdke and Stephan Günnemann and Leo Schwinn},
  year = {2026},
  month = jun,
  eprint = {2606.03647},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.03647}
}