June 2026Unreviewed
BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks
Ivan Sabolić, Marin Oršić, Josip Šarić, Sven Lončarić
Abstract
Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that this paradigm is highly vulnerable to backdoor attacks, and that existing defenses are ineffective in open-ended generation settings. In response, we propose BYORn, a backdoor-robust fine-tuning framework motivated by the observation that poisoned target responses are often semantically implausible given the corresponding image-text inputs and a pr
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{sabolic2026byorn,
title = {{BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks}},
author = {Ivan Sabolić and Marin Oršić and Josip Šarić and Sven Lončarić},
year = {2026},
month = jun,
eprint = {2606.02947},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.02947}
}