Skip to content
Search
paperJune 2026Unreviewed

BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks

Ivan Sabolić, Marin Oršić, Josip Šarić, Sven Lončarić

Abstract

Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that this paradigm is highly vulnerable to backdoor attacks, and that existing defenses are ineffective in open-ended generation settings. In response, we propose BYORn, a backdoor-robust fine-tuning framework motivated by the observation that poisoned target responses are often semantically implausible given the corresponding image-text inputs and a pr

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{sabolic2026byorn,
  title = {{BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks}},
  author = {Ivan Sabolić and Marin Oršić and Josip Šarić and Sven Lončarić},
  year = {2026},
  month = jun,
  eprint = {2606.02947},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.02947}
}