Skip to content
Search
paperFebruary 2024Unreviewed

Defending Jailbreak Prompts via In-Context Adversarial Game

Yujun Zhou, Yufei Han, Haomin Zhuang, Taicheng Guo, Kehan Guo, Zhenwen Liang, Hongyan Bao, Xiangliang Zhang

Conference on Empirical Methods in Natural Language Processing

Abstract

Large Language Models (LLMs) demonstrate remarkable capabilities across diverse applications. However, concerns regarding their security, particularly the vulnerability to jailbreak attacks, persist. Drawing inspiration from adversarial training in deep learning and LLM agent learning processes, we introduce the In-Context Adversarial Game (ICAG) for defending against jailbreaks without the need for fine-tuning. ICAG leverages agent learning to conduct an adversarial game, aiming to dynamically

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@inproceedings{zhou2024defending,
  title = {{Defending Jailbreak Prompts via In-Context Adversarial Game}},
  author = {Yujun Zhou and Yufei Han and Haomin Zhuang and Taicheng Guo and Kehan Guo and Zhenwen Liang and Hongyan Bao and Xiangliang Zhang},
  year = {2024},
  month = feb,
  booktitle = {Conference on Empirical Methods in Natural Language Processing},
  eprint = {2402.13148},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2402.13148},
  url = {https://www.semanticscholar.org/paper/50ceabc6aa41e08480fa5976342bfe04bb47bce3}
}