April 2026Unreviewed
Jailbreaking Frontier Foundation Models Through Intention Deception
Xinhe Wang, Katia Sycara, Yaqi Xie
Abstract
Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's intent. It has been found that this binary training regime often leads to brittleness, since the user intent cannot reliably be evaluated, especially if the attacker obfuscates their intent, and also makes the system seem unhelpful. In response, frontier models, s
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{wang2026jailbreaking,
title = {{Jailbreaking Frontier Foundation Models Through Intention Deception}},
author = {Xinhe Wang and Katia Sycara and Yaqi Xie},
year = {2026},
month = apr,
eprint = {2604.24082},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.24082}
}