August 2026Unreviewed
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
Zeyu Feng, Qingyuan Wu, Yu-Zhe Luo, Hua Cheng
Abstract
Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{feng2026psychjail,
title = {{PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies}},
author = {Zeyu Feng and Qingyuan Wu and Yu-Zhe Luo and Hua Cheng},
year = {2026},
month = aug,
eprint = {2608.23028},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/c8c62a6283d655086c1b928e4ec722888ff5b8fb}
}