Skip to content
Search
paperSeptember 2026Unreviewed

Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework

Shu-Yao Xiao, Sheng-Ling Wang, Hao-Yu Niu, Ke Chao, Changbo Xu, Xin-Ran Duan, Chao-Yong Jiang

Abstract

Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal little about how different information sources shape the generation process. We propose a causal and temporal evaluation framework that traces the evolving roles of visual input, question text, and generated prefixes during autoregressive decoding. Grounded in a Structural Causal Model, we use interventions and backdoor

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{xiao2026who,
  title = {{Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework}},
  author = {Shu-Yao Xiao and Sheng-Ling Wang and Hao-Yu Niu and Ke Chao and Changbo Xu and Xin-Ran Duan and Chao-Yong Jiang},
  year = {2026},
  month = sep,
  eprint = {2609.02000},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/c1f5686528f9fe5fd5b7db14e9d275a6a17ac9ab}
}