August 2026Unreviewed
Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models
Yusuke Hirota, Michael Boone, Arun George Zachariah, J. Varghese, Y. Wang, Boyi Li, Ryo Hachiuma
Abstract
We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely on prompts that ask models to infer attributes of people in images (e.g.,"Is this person a CEO or a secretary?"). However, we find that LVLMs with strong guardrails, such as GPT and Claude, often refuse these prompts, making evaluations unreliable. To address this, we change the prior evaluation paradigm by decoupling the task from the depicted pe
Categories
Cite
@misc{hirota2026guardrailagnostic,
title = {{Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models}},
author = {Yusuke Hirota and Michael Boone and Arun George Zachariah and J. Varghese and Y. Wang and Boyi Li and Ryo Hachiuma},
year = {2026},
month = aug,
eprint = {2608.29590},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/22de4aaa0b81f2b475d9aa1ee401cc4f5c46172c}
}