Skip to content
Search
paperApril 2026Unreviewed

Jailbreaking Frontier Foundation Models Through Intention Deception

Xinhe Wang, Katia Sycara, Yaqi Xie

Abstract

Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's intent. It has been found that this binary training regime often leads to brittleness, since the user intent cannot reliably be evaluated, especially if the attacker obfuscates their intent, and also makes the system seem unhelpful. In response, frontier models, s

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{wang2026jailbreaking,
  title = {{Jailbreaking Frontier Foundation Models Through Intention Deception}},
  author = {Xinhe Wang and Katia Sycara and Yaqi Xie},
  year = {2026},
  month = apr,
  eprint = {2604.24082},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.24082}
}