Skip to content
Search
paperJune 2024Unreviewed

Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs

Fan Liu, Zhao Xu, Hao Liu

arXiv.org

Abstract

Although safely enhanced Large Language Models (LLMs) have achieved remarkable success in tackling various complex tasks in a zero-shot manner, they remain susceptible to jailbreak attacks, particularly the unknown jailbreak attack. To enhance LLMs' generalized defense capabilities, we propose a two-stage adversarial tuning framework, which generates adversarial prompts to explore worst-case scenarios by optimizing datasets containing pairs of adversarial prompts and their safe responses. In the

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{liu2024adversarial,
  title = {{Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs}},
  author = {Fan Liu and Zhao Xu and Hao Liu},
  year = {2024},
  month = jun,
  eprint = {2406.06622},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2406.06622},
  url = {https://www.semanticscholar.org/paper/9223a64feb573e62498c2ca914ed97557c580167}
}