September 2026Unreviewed
Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails
Suyoung Lee, Myungsub Choi
Abstract
Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instruction-screenshot pairs with instruction-relative labels, pixel-exact evidence boxes, and matched image-side counterfactuals. Across six VLMs, two models with nearly identical aver
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{lee2026beyond,
title = {{Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails}},
author = {Suyoung Lee and Myungsub Choi},
year = {2026},
month = sep,
eprint = {2609.05535},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/1777bccb47960930d9560230d9b519fed1ab95ca}
}