September 2026Unreviewed
CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models
Fei-Fei Liu, Jintao Cheng, Chi-Man Vong, Xiaoyu Tang
Abstract
Training-free collaborative pipelines that integrate Vision Foundation Models such as CLIP, SAM, and DINO achieve strong open-vocabulary dense prediction and are increasingly deployed in safety-critical applications. The security of these systems is commonly assumed to follow from the robustness of their individual models. We challenge this assumption. We identify a vulnerability shared by every collaborative pipeline: each model consumes the intermediate output of another without verifying sema
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{liu2026crack,
title = {{CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models}},
author = {Fei-Fei Liu and Jintao Cheng and Chi-Man Vong and Xiaoyu Tang},
year = {2026},
month = sep,
eprint = {2609.07499},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/c6f9d684d9b92f5baf1af32a027c1320fb953039}
}