Skip to content
Search
paper2024ReviewedOpen access

GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, Zhaopeng Tu

ICLR 2024

Abstract

Demonstrates that LLMs can be jailbroken using cipher-based encoding, bypassing safety training designed for natural language.

Categories

#cipher#encoding#bypass

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Cite

@inproceedings{yuan2024gpt4,
  title = {{GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher}},
  author = {Youliang Yuan and Wenxiang Jiao and Wenxuan Wang and Jen-tse Huang and Pinjia He and Shuming Shi and Zhaopeng Tu},
  year = {2024},
  booktitle = {ICLR 2024},
  eprint = {2308.06463},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2308.06463}
}