Skip to content
Search
paperAugust 2026Unreviewed

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt, Simone Ponzetto

Abstract

Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{klerings2026mood,
  title = {{Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment}},
  author = {Alina Klerings and Jannik Brinkmann and Heiner Stuckenschmidt and Simone Ponzetto},
  year = {2026},
  month = aug,
  eprint = {2608.05409},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/d6d288940807fdef4800e0acdb1b2c14e657baed}
}