August 2026Unreviewed
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
Xinzhe Huang, Biwu Yao, Kedong Xiu, Mengnan Zhao, Di Wang, Puning Zhao, Tianhang Zheng
Abstract
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete token sequence to a discrete safety label. However, this paradigm has two limitations: First, safety assessment is inherently an uncertain problem, particularly during the early generation state. Second, relying solely on discrete token sequences discards the rich pro
Categories
Cite
@misc{huang2026probguard,
title = {{ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions}},
author = {Xinzhe Huang and Biwu Yao and Kedong Xiu and Mengnan Zhao and Di Wang and Puning Zhao and Tianhang Zheng},
year = {2026},
month = aug,
eprint = {2608.10621},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/f6c0d06532a9084259bdd4e94343c9a1e98e4f23}
}