August 2026Unreviewed
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
Feitong Qiao, Liren Peng, Shiming Ren, Aishwarya Jadhav, Arghavan Bahadorinejad, Marinette Chen, Muhan Zhang, Abdulaziz Suria, Gennevi Lu, Anish Das Sarma
Abstract
Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as a generation problem: produce attacks that break the model. We argue it is better framed as a search problem: discover, organize, and iteratively refine a diverse archive of attack strategies, producing a structured map
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{qiao2026evoflint,
title = {{EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities}},
author = {Feitong Qiao and Liren Peng and Shiming Ren and Aishwarya Jadhav and Arghavan Bahadorinejad and Marinette Chen and Muhan Zhang and Abdulaziz Suria and Gennevi Lu and Anish Das Sarma},
year = {2026},
month = aug,
eprint = {2609.00487},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.00487}
}