Skip to content
Search
paperAugust 2026Unreviewed

Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning

Yuyang Luo, Kai Shu

Abstract

Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains where fairness is a critical safety requirement. In practice, these models are continually updated through continual learning (CL) to adapt to evolving tasks and data distributions. Prior work has shown that backdoor attacks can manipulate MLLM responses through hidden triggers, but naively implanted backdoors degrade as models undergo subsequent updates of CL. Although fairness has emerged as a central con

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{luo2026anchoring,
  title = {{Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning}},
  author = {Yuyang Luo and Kai Shu},
  year = {2026},
  month = aug,
  eprint = {2608.21577},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/8b1455c5a420714262b830d1af702fd2478832c6}
}