Skip to content
Search
paperMay 2026Unreviewed

DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models

Ye Sun, Xin Wang, Jiaming Zhang, Yifeng Gao, Yixu Wang, Yifan Ding, Qixian Zhang, Henghui Ding, Xingjun Ma, Yu-Gang Jiang

Abstract

While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, traditional adversarial attacks are typically limited to single, predefined objectives, tightly coupling each attack to a specific model or task, which restricts their scalability and flexibility in real-world scenarios. In this work, we present DarkLLM, a novel attack framework that trains an LLM to translate natural-language

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{sun2026darkllm,
  title = {{DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models}},
  author = {Ye Sun and Xin Wang and Jiaming Zhang and Yifeng Gao and Yixu Wang and Yifan Ding and Qixian Zhang and Henghui Ding and Xingjun Ma and Yu-Gang Jiang},
  year = {2026},
  month = may,
  eprint = {2605.18868},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.18868}
}