Skip to content
Search
paperJune 2026Unreviewed

RogueMerge: Robust and Unified Attacks against LLM Model Merging

Jinghuai Zhang, Yetian He, Kunlin Cai, Han Zhao, Fnu Suya, Yuan Tian

Abstract

Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using stat

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{zhang2026roguemerge,
  title = {{RogueMerge: Robust and Unified Attacks against LLM Model Merging}},
  author = {Jinghuai Zhang and Yetian He and Kunlin Cai and Han Zhao and Fnu Suya and Yuan Tian},
  year = {2026},
  month = jun,
  eprint = {2606.03344},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.03344}
}