June 2026Unreviewed
RogueMerge: Robust and Unified Attacks against LLM Model Merging
Jinghuai Zhang, Yetian He, Kunlin Cai, Han Zhao, Fnu Suya, Yuan Tian
Abstract
Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using stat
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{zhang2026roguemerge,
title = {{RogueMerge: Robust and Unified Attacks against LLM Model Merging}},
author = {Jinghuai Zhang and Yetian He and Kunlin Cai and Han Zhao and Fnu Suya and Yuan Tian},
year = {2026},
month = jun,
eprint = {2606.03344},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.03344}
}