April 2026Unreviewed
Jailbreaking Large Language Models with Morality Attacks
Ying Su, Mingen Zheng, Weili Diao, Haoran Li
Abstract
Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Research towards pluralism alignment has many efforts in enhancing the learning of large language models (LLMs) to accomplish pluralism. Although this is essential, the robustness of LLMs to produce moral content over pluralistic values is still under exploration.Inspired by the astonishing persuasion abilities via jailbreak prompts, we propose to lev
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{su2026jailbreaking,
title = {{Jailbreaking Large Language Models with Morality Attacks}},
author = {Ying Su and Mingen Zheng and Weili Diao and Haoran Li},
year = {2026},
month = apr,
eprint = {2604.17053},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.17053}
}