March 2026Unreviewed
Test-Time Attention Purification for Backdoored Large Vision Language Models
Zhifang Zhang, Bojun Yang, Shuo He, Weitong Chen, Wei Emma Zhang, Olaf Maennel, Lei Feng, Miao Xu
Abstract
Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded samples into the training data to implant behaviors that can be maliciously activated at test time. Existing defenses typically rely on retraining backdoored parameters (e.g., adapters or LoRA modules) with clean data, which is computationally expensive and often degrades model performance. In this work, we provide a new
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{zhang2026testtime,
title = {{Test-Time Attention Purification for Backdoored Large Vision Language Models}},
author = {Zhifang Zhang and Bojun Yang and Shuo He and Weitong Chen and Wei Emma Zhang and Olaf Maennel and Lei Feng and Miao Xu},
year = {2026},
month = mar,
eprint = {2603.12989},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2603.12989}
}