Skip to content
Search
paperAugust 2026Unreviewed

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

Zeyu Feng, Qingyuan Wu, Yu-Zhe Luo, Hua Cheng

Abstract

Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{feng2026psychjail,
  title = {{PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies}},
  author = {Zeyu Feng and Qingyuan Wu and Yu-Zhe Luo and Hua Cheng},
  year = {2026},
  month = aug,
  eprint = {2608.23028},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/c8c62a6283d655086c1b928e4ec722888ff5b8fb}
}