May 2026Unreviewed
Laundering AI Authority with Adversarial Examples
Jie Zhang, Pura Peetathawatchai, Florian Tramèr, Avital Shafran
Abstract
Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial examples break this assumption, enabling \emph{AI authority laundering}: an attacker subtly perturbs an image so that the VLM produces confident and authoritative responses about the \emph{wrong} input. Unlike jailbreaks or
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhang2026laundering,
title = {{Laundering AI Authority with Adversarial Examples}},
author = {Jie Zhang and Pura Peetathawatchai and Florian Tramèr and Avital Shafran},
year = {2026},
month = may,
eprint = {2605.04261},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.04261}
}