Skip to content
Search
paperAugust 2026Unreviewed

Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models

Yusuke Hirota, Michael Boone, Arun George Zachariah, J. Varghese, Y. Wang, Boyi Li, Ryo Hachiuma

Abstract

We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely on prompts that ask models to infer attributes of people in images (e.g.,"Is this person a CEO or a secretary?"). However, we find that LVLMs with strong guardrails, such as GPT and Claude, often refuse these prompts, making evaluations unreliable. To address this, we change the prior evaluation paradigm by decoupling the task from the depicted pe

Categories

Cite

@misc{hirota2026guardrailagnostic,
  title = {{Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models}},
  author = {Yusuke Hirota and Michael Boone and Arun George Zachariah and J. Varghese and Y. Wang and Boyi Li and Ryo Hachiuma},
  year = {2026},
  month = aug,
  eprint = {2608.29590},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/22de4aaa0b81f2b475d9aa1ee401cc4f5c46172c}
}