May 2026Unreviewed
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
Ye Sun, Xin Wang, Jiaming Zhang, Yifeng Gao, Yixu Wang, Yifan Ding, Qixian Zhang, Henghui Ding, Xingjun Ma, Yu-Gang Jiang
Abstract
While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, traditional adversarial attacks are typically limited to single, predefined objectives, tightly coupling each attack to a specific model or task, which restricts their scalability and flexibility in real-world scenarios. In this work, we present DarkLLM, a novel attack framework that trains an LLM to translate natural-language
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{sun2026darkllm,
title = {{DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models}},
author = {Ye Sun and Xin Wang and Jiaming Zhang and Yifeng Gao and Yixu Wang and Yifan Ding and Qixian Zhang and Henghui Ding and Xingjun Ma and Yu-Gang Jiang},
year = {2026},
month = may,
eprint = {2605.18868},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.18868}
}