Skip to content
Search
paperMay 2024Unreviewed

WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response

Tianrong Zhang, Bochuan Cao, Yuanpu Cao, Lu Lin, Prasenjit Mitra, Jinghui Chen

North American Chapter of the Association for Computational Linguistics

Abstract

The recent breakthrough in large language models (LLMs) such as ChatGPT has revolutionized production processes at an unprecedented pace. Alongside this progress also comes mounting concerns about LLMs' susceptibility to jailbreaking attacks, which leads to the generation of harmful or unsafe content. While safety alignment measures have been implemented in LLMs to mitigate existing jailbreak attempts and force them to become increasingly complicated, it is still far from perfect. In this paper,

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@article{zhang2024wordgame,
  title = {{WordGame: Efficient \& Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response}},
  author = {Tianrong Zhang and Bochuan Cao and Yuanpu Cao and Lu Lin and Prasenjit Mitra and Jinghui Chen},
  year = {2024},
  month = may,
  journal = {North American Chapter of the Association for Computational Linguistics},
  eprint = {2405.14023},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2405.14023},
  url = {https://www.semanticscholar.org/paper/8db6ff37617c5d3a6aec9e40e5e829a735d0c0cf}
}