May 2024Unreviewed
WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response
Tianrong Zhang, Bochuan Cao, Yuanpu Cao, Lu Lin, Prasenjit Mitra, Jinghui Chen
North American Chapter of the Association for Computational Linguistics
Abstract
The recent breakthrough in large language models (LLMs) such as ChatGPT has revolutionized production processes at an unprecedented pace. Alongside this progress also comes mounting concerns about LLMs' susceptibility to jailbreaking attacks, which leads to the generation of harmful or unsafe content. While safety alignment measures have been implemented in LLMs to mitigate existing jailbreak attempts and force them to become increasingly complicated, it is still far from perfect. In this paper,
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{zhang2024wordgame,
title = {{WordGame: Efficient \& Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response}},
author = {Tianrong Zhang and Bochuan Cao and Yuanpu Cao and Lu Lin and Prasenjit Mitra and Jinghui Chen},
year = {2024},
month = may,
journal = {North American Chapter of the Association for Computational Linguistics},
eprint = {2405.14023},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2405.14023},
url = {https://www.semanticscholar.org/paper/8db6ff37617c5d3a6aec9e40e5e829a735d0c0cf}
}