July 2026Unreviewed
HeteroGuard: Constructing Safety Boundaries for Large Language Models via Heterogeneous Weak-Model Collaboration
Ping Chen, Yin Cai, Shuting Zhang, Yi Wang, Xingqiu Shen, Yuting Shang, Hong Zou
Security and Safety
Abstract
Large language models (LLMs) face safety risks in deployment, but existing evaluation methods rely on single powerful judges and treat safety as binary classification. This paper proposes HeteroGuard, a dynamic heterogeneous redundancy (DHR)-inspired framework that constructs a safety boundary around LLM outputs through heterogeneous weak-model collaboration. HeteroGuard deploys three diverse weak judges and aggregates their decisions via majority voting to define a boundary that organizes respo
Categories
Cite
@article{chen2026heteroguard,
title = {{HeteroGuard: Constructing Safety Boundaries for Large Language Models via Heterogeneous Weak-Model Collaboration}},
author = {Ping Chen and Yin Cai and Shuting Zhang and Yi Wang and Xingqiu Shen and Yuting Shang and Hong Zou},
year = {2026},
month = jul,
journal = {Security and Safety},
doi = {10.1051/sands/2026018},
url = {https://doi.org/10.1051/sands/2026018}
}