Skip to content
Search
paperJune 2026Unreviewed

Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks

Haoming Wen, Shi Chen, Qingyu Shi, Siyuan Liu, Minrui Luo, Jingzhao Zhang, Tianxing He

Abstract

Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised finetuning (SFT) on poisoned datasets. Existing alignment-stage defenses are primarily designed to defend against attacks that use parameter-efficient finetuning methods. However, they fail to defend against stronger attacks that use full-parameter finetuning. In this paper, we propose Patcher, a method inspired by adv

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{wen2026defending,
  title = {{Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks}},
  author = {Haoming Wen and Shi Chen and Qingyu Shi and Siyuan Liu and Minrui Luo and Jingzhao Zhang and Tianxing He},
  year = {2026},
  month = jun,
  eprint = {2606.07970},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.07970}
}