Skip to content
Search
paperAugust 2026Unreviewed

Validity-Aware Jailbreak Evaluation for Large Language Models

Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran

Abstract

Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather than correctness. We identify a key limitation in existing evaluations: many jailbreak intents depend on instructional validity rather than epistemic factuality, allowing realistic-looking responses to be labeled successful despite being factually

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{wu2026validityaware,
  title = {{Validity-Aware Jailbreak Evaluation for Large Language Models}},
  author = {Qilong Wu and Sahil Wadhwa and Pranab Mohanty and Giri Iyengar and Varun Chandrasekaran},
  year = {2026},
  month = aug,
  eprint = {2609.00498},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/08abe34670f9011fc3ff59cbab6c6328b2480af4}
}