Skip to content
Search
paperJuly 2024Unreviewed

JailbreakHunter: A Visual Analytics Approach for Jailbreak Prompts Discovery From Large-Scale Human-LLM Conversational Datasets

Zhihua Jin, Shiyi Liu, Haotian Li, Xun Zhao, Huamin Qu

IEEE Transactions on Visualization and Computer Graphics

Abstract

Large Language Models (LLMs) have gained significant attention but also raised concerns due to the risk of misuse. Jailbreak prompts, a popular type of adversarial attack towards LLMs, have appeared and constantly evolved to breach the safety protocols of LLMs. To address this issue, LLMs are regularly updated with safety patches based on reported jailbreak prompts. However, malicious users often keep their successful jailbreak prompts private to exploit LLMs. To uncover these private jailbreak

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@article{jin2024jailbreakhunter,
  title = {{JailbreakHunter: A Visual Analytics Approach for Jailbreak Prompts Discovery From Large-Scale Human-LLM Conversational Datasets}},
  author = {Zhihua Jin and Shiyi Liu and Haotian Li and Xun Zhao and Huamin Qu},
  year = {2024},
  month = jul,
  journal = {IEEE Transactions on Visualization and Computer Graphics},
  eprint = {2407.03045},
  archivePrefix = {arXiv},
  doi = {10.1109/TVCG.2025.3557568},
  url = {https://www.semanticscholar.org/paper/b679c8166de2332184d22e988e269739778f03a0}
}