Skip to content
Search
paperAugust 2026Unreviewed

Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs

Beining Xu, Hairui Wang, Jiaxin Wang, Changsheng Chen, Anirban Chakraborty

Abstract

While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding MLLMs for identity document processing, this paper investigates the privacy issues inherent in Key Information Extraction (KIE) tasks. We reveal that when input images lack sufficient visual evidence, these models often rely on memorized field relations from training data to infer

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0024.000Infer Training Data Membership

Suggested from the entry's categories.

Cite

@misc{xu2026beyond,
  title = {{Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs}},
  author = {Beining Xu and Hairui Wang and Jiaxin Wang and Changsheng Chen and Anirban Chakraborty},
  year = {2026},
  month = aug,
  eprint = {2608.12911},
  archivePrefix = {arXiv},
  doi = {10.1145/3767308.3835901},
  url = {https://www.semanticscholar.org/paper/8b616f7b4ce3fe4934e0fd0beb28e9f6ca206b71}
}