Skip to content
Search
paperMay 2026Unreviewed

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

Xiaoyu Lu, Xianglin Yang, Haijun Liu, Jiahao Liu, Kuntai Cai, Yan Xiao, J. Dong

Abstract

The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on constructed benchmarks to assess safety from predefined perspectives, or employ dynamic red-teaming to probe potential vulnerabilities. While effective, these approaches face challenges, as they depend heavily on expert domain knowledge, offer limited systematic guarantees, and are vulnerable to rapid obsolescence. To address these limitations, we i

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{lu2026inverting,
  title = {{Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications}},
  author = {Xiaoyu Lu and Xianglin Yang and Haijun Liu and Jiahao Liu and Kuntai Cai and Yan Xiao and J. Dong},
  year = {2026},
  month = may,
  eprint = {2605.24883},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/6a1f7108e15400c9dd3c2054125afaffebd03299}
}