Skip to content
Search
paperAugust 2026Unreviewed

Inverting the Hidden: Unveiling Multimodal Privacy Leakage in Collaborative LVLM Inference

Shuaifan Jin, Zhibo Wang, Qiyuan Wang, Yiting Han, Yajie Zhou, Yuanfan Zhang, Jiahui Hu, Xiaoyi Pang

Abstract

Collaborative inference deploys Large Vision-Language Models (LVLMs) by partitioning computation between edge devices and the cloud. While withholding raw inputs supposedly ensures privacy, transmitting intermediate hidden states exposes a critical attack surface. However, it remains unclear whether deep-layer LVLM hidden states retain recoverable private information, given that visual content has been projected into the language embedding space. To address this concern, we theoretically analyze

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0024.000Infer Training Data Membership

Suggested from the entry's categories.

Cite

@misc{jin2026inverting,
  title = {{Inverting the Hidden: Unveiling Multimodal Privacy Leakage in Collaborative LVLM Inference}},
  author = {Shuaifan Jin and Zhibo Wang and Qiyuan Wang and Yiting Han and Yajie Zhou and Yuanfan Zhang and Jiahui Hu and Xiaoyi Pang},
  year = {2026},
  month = aug,
  eprint = {2608.01020},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/887f61b15e19ba216107ce815c972613e659e020}
}