Skip to content
Search
paperSeptember 2026Unreviewed

Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails

Suyoung Lee, Myungsub Choi

Abstract

Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instruction-screenshot pairs with instruction-relative labels, pixel-exact evidence boxes, and matched image-side counterfactuals. Across six VLMs, two models with nearly identical aver

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{lee2026beyond,
  title = {{Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails}},
  author = {Suyoung Lee and Myungsub Choi},
  year = {2026},
  month = sep,
  eprint = {2609.05535},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/1777bccb47960930d9560230d9b519fed1ab95ca}
}