July 2024Unreviewed
JailbreakHunter: A Visual Analytics Approach for Jailbreak Prompts Discovery From Large-Scale Human-LLM Conversational Datasets
Zhihua Jin, Shiyi Liu, Haotian Li, Xun Zhao, Huamin Qu
IEEE Transactions on Visualization and Computer Graphics
Abstract
Large Language Models (LLMs) have gained significant attention but also raised concerns due to the risk of misuse. Jailbreak prompts, a popular type of adversarial attack towards LLMs, have appeared and constantly evolved to breach the safety protocols of LLMs. To address this issue, LLMs are regularly updated with safety patches based on reported jailbreak prompts. However, malicious users often keep their successful jailbreak prompts private to exploit LLMs. To uncover these private jailbreak
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{jin2024jailbreakhunter,
title = {{JailbreakHunter: A Visual Analytics Approach for Jailbreak Prompts Discovery From Large-Scale Human-LLM Conversational Datasets}},
author = {Zhihua Jin and Shiyi Liu and Haotian Li and Xun Zhao and Huamin Qu},
year = {2024},
month = jul,
journal = {IEEE Transactions on Visualization and Computer Graphics},
eprint = {2407.03045},
archivePrefix = {arXiv},
doi = {10.1109/TVCG.2025.3557568},
url = {https://www.semanticscholar.org/paper/b679c8166de2332184d22e988e269739778f03a0}
}