April 2026Unreviewed
Adaptive Instruction Composition for Automated LLM Red-Teaming
Jesse Zymet, Andy Luo, Swapnil Shinde, Sahil Wadhwa, Emily Chen
Abstract
Many approaches to LLM red-teaming leverage an attacker LLM to discover jailbreaks against a target. Several of them task the attacker with identifying effective strategies through trial and error, resulting in a semantically limited range of successes. Another approach discovers diverse attacks by combining crowdsourced harmful queries and tactics into instructions for the attacker, but does so at random, limiting effectiveness. This article introduces a novel framework, Adaptive Instruction Co
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{zymet2026adaptive,
title = {{Adaptive Instruction Composition for Automated LLM Red-Teaming}},
author = {Jesse Zymet and Andy Luo and Swapnil Shinde and Sahil Wadhwa and Emily Chen},
year = {2026},
month = apr,
eprint = {2604.21159},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.21159}
}