Skip to content
Search
paperMay 2026Unreviewed

Laundering AI Authority with Adversarial Examples

Jie Zhang, Pura Peetathawatchai, Florian Tramèr, Avital Shafran

Abstract

Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial examples break this assumption, enabling \emph{AI authority laundering}: an attacker subtly perturbs an image so that the VLM produces confident and authoritative responses about the \emph{wrong} input. Unlike jailbreaks or

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026laundering,
  title = {{Laundering AI Authority with Adversarial Examples}},
  author = {Jie Zhang and Pura Peetathawatchai and Florian Tramèr and Avital Shafran},
  year = {2026},
  month = may,
  eprint = {2605.04261},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.04261}
}