Skip to content
Search
paperSeptember 2026Unreviewed

Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges

Rui Yang, Shuang Huang, Junhua Liu, Ziqi Zhao, Qingzhong Yan, Yuhang Sun, Cong Liu, Guoping Hu, Rui Mei, Jing Shao

Abstract

Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where linguistic variation and adversarial transformations can obscure risky intent. We introduce C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation. It comprises 538 base queries and 8,877 adversarial queries answered by f

Categories

Cite

@misc{yang2026who,
  title = {{Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges}},
  author = {Rui Yang and Shuang Huang and Junhua Liu and Ziqi Zhao and Qingzhong Yan and Yuhang Sun and Cong Liu and Guoping Hu and Rui Mei and Jing Shao},
  year = {2026},
  month = sep,
  eprint = {2609.01210},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.01210}
}