Skip to content
Search
paperMarch 2026Unreviewed

Test-Time Attention Purification for Backdoored Large Vision Language Models

Zhifang Zhang, Bojun Yang, Shuo He, Weitong Chen, Wei Emma Zhang, Olaf Maennel, Lei Feng, Miao Xu

Abstract

Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded samples into the training data to implant behaviors that can be maliciously activated at test time. Existing defenses typically rely on retraining backdoored parameters (e.g., adapters or LoRA modules) with clean data, which is computationally expensive and often degrades model performance. In this work, we provide a new

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{zhang2026testtime,
  title = {{Test-Time Attention Purification for Backdoored Large Vision Language Models}},
  author = {Zhifang Zhang and Bojun Yang and Shuo He and Weitong Chen and Wei Emma Zhang and Olaf Maennel and Lei Feng and Miao Xu},
  year = {2026},
  month = mar,
  eprint = {2603.12989},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2603.12989}
}