February 2025Unreviewed
Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
Zixuan Weng, Xiaolong Jin, Jinyuan Jia, Xiangyu Zhang
Conference on Empirical Methods in Natural Language Processing
Abstract
Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A key challenge is jailbreak, where adversarial prompts bypass built-in safeguards to elicit harmful disallowed outputs. Inspired by psychological foot-in-the-door principles, we introduce FITD,a novel multi-turn jailbreak method that leverages the phenomenon where minor initial commitments lower resistance to more significant or more unethical transgressions. Our approach progress
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@inproceedings{weng2025footinthedoor,
title = {{Foot-In-The-Door: A Multi-turn Jailbreak for LLMs}},
author = {Zixuan Weng and Xiaolong Jin and Jinyuan Jia and Xiangyu Zhang},
year = {2025},
month = feb,
booktitle = {Conference on Empirical Methods in Natural Language Processing},
eprint = {2502.19820},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2502.19820},
url = {https://www.semanticscholar.org/paper/5b79112817c0115d3db49245311982b623270422}
}