May 2026Unreviewed
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
Junke Zhang, Jianwei Wang, Sishuo Chen, Yizhang He, Qingshuai Feng, Zhengyi Yang
Abstract
Jailbreak attacks on large language models (LLMs) aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is especially important for safety evaluation, where the attacker observes only model outputs and needs to automatically search for effective adversarial prompts. Existing black-box jailbreak methods either depend on sample-wise heuristic search or leverage attack experience through accumulating strategy pools or method libraries, lack
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhang2026evolving,
title = {{Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking}},
author = {Junke Zhang and Jianwei Wang and Sishuo Chen and Yizhang He and Qingshuai Feng and Zhengyi Yang},
year = {2026},
month = may,
eprint = {2605.29237},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.29237}
}