June 2026Unreviewed
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
Haoming Wen, Shi Chen, Qingyu Shi, Siyuan Liu, Minrui Luo, Jingzhao Zhang, Tianxing He
Abstract
Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised finetuning (SFT) on poisoned datasets. Existing alignment-stage defenses are primarily designed to defend against attacks that use parameter-efficient finetuning methods. However, they fail to defend against stronger attacks that use full-parameter finetuning. In this paper, we propose Patcher, a method inspired by adv
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{wen2026defending,
title = {{Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks}},
author = {Haoming Wen and Shi Chen and Qingyu Shi and Siyuan Liu and Minrui Luo and Jingzhao Zhang and Tianxing He},
year = {2026},
month = jun,
eprint = {2606.07970},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.07970}
}