August 2026Unreviewed
Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning
Yuyang Luo, Kai Shu
Abstract
Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains where fairness is a critical safety requirement. In practice, these models are continually updated through continual learning (CL) to adapt to evolving tasks and data distributions. Prior work has shown that backdoor attacks can manipulate MLLM responses through hidden triggers, but naively implanted backdoors degrade as models undergo subsequent updates of CL. Although fairness has emerged as a central con
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{luo2026anchoring,
title = {{Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning}},
author = {Yuyang Luo and Kai Shu},
year = {2026},
month = aug,
eprint = {2608.21577},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/8b1455c5a420714262b830d1af702fd2478832c6}
}