Skip to content
Search
paperApril 2026Unreviewed

Jailbreaking Large Language Models with Morality Attacks

Ying Su, Mingen Zheng, Weili Diao, Haoran Li

Abstract

Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Research towards pluralism alignment has many efforts in enhancing the learning of large language models (LLMs) to accomplish pluralism. Although this is essential, the robustness of LLMs to produce moral content over pluralistic values is still under exploration.Inspired by the astonishing persuasion abilities via jailbreak prompts, we propose to lev

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{su2026jailbreaking,
  title = {{Jailbreaking Large Language Models with Morality Attacks}},
  author = {Ying Su and Mingen Zheng and Weili Diao and Haoran Li},
  year = {2026},
  month = apr,
  eprint = {2604.17053},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.17053}
}