May 2026Unreviewed
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
Xiaoyu Lu, Xianglin Yang, Haijun Liu, Jiahao Liu, Kuntai Cai, Yan Xiao, J. Dong
Abstract
The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on constructed benchmarks to assess safety from predefined perspectives, or employ dynamic red-teaming to probe potential vulnerabilities. While effective, these approaches face challenges, as they depend heavily on expert domain knowledge, offer limited systematic guarantees, and are vulnerable to rapid obsolescence. To address these limitations, we i
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{lu2026inverting,
title = {{Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications}},
author = {Xiaoyu Lu and Xianglin Yang and Haijun Liu and Jiahao Liu and Kuntai Cai and Yan Xiao and J. Dong},
year = {2026},
month = may,
eprint = {2605.24883},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/6a1f7108e15400c9dd3c2054125afaffebd03299}
}