June 2024Unreviewed
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
Fan Liu, Zhao Xu, Hao Liu
arXiv.org
Abstract
Although safely enhanced Large Language Models (LLMs) have achieved remarkable success in tackling various complex tasks in a zero-shot manner, they remain susceptible to jailbreak attacks, particularly the unknown jailbreak attack. To enhance LLMs' generalized defense capabilities, we propose a two-stage adversarial tuning framework, which generates adversarial prompts to explore worst-case scenarios by optimizing datasets containing pairs of adversarial prompts and their safe responses. In the
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{liu2024adversarial,
title = {{Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs}},
author = {Fan Liu and Zhao Xu and Hao Liu},
year = {2024},
month = jun,
eprint = {2406.06622},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2406.06622},
url = {https://www.semanticscholar.org/paper/9223a64feb573e62498c2ca914ed97557c580167}
}