Skip to content
Search
paperSeptember 2026Unreviewed

CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models

Fei-Fei Liu, Jintao Cheng, Chi-Man Vong, Xiaoyu Tang

Abstract

Training-free collaborative pipelines that integrate Vision Foundation Models such as CLIP, SAM, and DINO achieve strong open-vocabulary dense prediction and are increasingly deployed in safety-critical applications. The security of these systems is commonly assumed to follow from the robustness of their individual models. We challenge this assumption. We identify a vulnerability shared by every collaborative pipeline: each model consumes the intermediate output of another without verifying sema

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{liu2026crack,
  title = {{CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models}},
  author = {Fei-Fei Liu and Jintao Cheng and Chi-Man Vong and Xiaoyu Tang},
  year = {2026},
  month = sep,
  eprint = {2609.07499},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/c6f9d684d9b92f5baf1af32a027c1320fb953039}
}