← Back to search
paper llmsec-2026-00120

Laundering AI Authority with Adversarial Examples

Jie Zhang, Pura Peetathawatchai, Florian Tramèr, Avital Shafran

2026-05

Abstract

Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial examples break this assumption, enabling \emph{AI authority laundering}: an attacker subtly perturbs an image so that the VLM produces confident and authoritative responses about the \emph{wrong} input. Unlike jailbreaks or

Cite This Resource

@article{llmsec202600120,
  title = {Laundering AI Authority with Adversarial Examples},
  author = {Jie Zhang and Pura Peetathawatchai and Florian Tramèr and Avital Shafran},
  year = {2026},
  url = {https://arxiv.org/abs/2605.04261},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.04261