Skip to content
Search
paperJuly 2026Unreviewed

HeteroGuard: Constructing Safety Boundaries for Large Language Models via Heterogeneous Weak-Model Collaboration

Ping Chen, Yin Cai, Shuting Zhang, Yi Wang, Xingqiu Shen, Yuting Shang, Hong Zou

Security and Safety

Abstract

Large language models (LLMs) face safety risks in deployment, but existing evaluation methods rely on single powerful judges and treat safety as binary classification. This paper proposes HeteroGuard, a dynamic heterogeneous redundancy (DHR)-inspired framework that constructs a safety boundary around LLM outputs through heterogeneous weak-model collaboration. HeteroGuard deploys three diverse weak judges and aggregates their decisions via majority voting to define a boundary that organizes respo

Categories

Cite

@article{chen2026heteroguard,
  title = {{HeteroGuard: Constructing Safety Boundaries for Large Language Models via Heterogeneous Weak-Model Collaboration}},
  author = {Ping Chen and Yin Cai and Shuting Zhang and Yi Wang and Xingqiu Shen and Yuting Shang and Hong Zou},
  year = {2026},
  month = jul,
  journal = {Security and Safety},
  doi = {10.1051/sands/2026018},
  url = {https://doi.org/10.1051/sands/2026018}
}