August 2026Unreviewed
Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt, Simone Ponzetto
Abstract
Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{klerings2026mood,
title = {{Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment}},
author = {Alina Klerings and Jannik Brinkmann and Heiner Stuckenschmidt and Simone Ponzetto},
year = {2026},
month = aug,
eprint = {2608.05409},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/d6d288940807fdef4800e0acdb1b2c14e657baed}
}