August 2026Unreviewed
Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
Beining Xu, Hairui Wang, Jiaxin Wang, Changsheng Chen, Anirban Chakraborty
Abstract
While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding MLLMs for identity document processing, this paper investigates the privacy issues inherent in Key Information Extraction (KIE) tasks. We reveal that when input images lack sufficient visual evidence, these models often rely on memorized field relations from training data to infer
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0024.000Infer Training Data Membership
Suggested from the entry's categories.
Cite
@misc{xu2026beyond,
title = {{Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs}},
author = {Beining Xu and Hairui Wang and Jiaxin Wang and Changsheng Chen and Anirban Chakraborty},
year = {2026},
month = aug,
eprint = {2608.12911},
archivePrefix = {arXiv},
doi = {10.1145/3767308.3835901},
url = {https://www.semanticscholar.org/paper/8b616f7b4ce3fe4934e0fd0beb28e9f6ca206b71}
}