Skip to content
Search
paperAugust 2026Unreviewed

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

Xinzhe Huang, Biwu Yao, Kedong Xiu, Mengnan Zhao, Di Wang, Puning Zhao, Tianhang Zheng

Abstract

Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete token sequence to a discrete safety label. However, this paradigm has two limitations: First, safety assessment is inherently an uncertain problem, particularly during the early generation state. Second, relying solely on discrete token sequences discards the rich pro

Categories

Cite

@misc{huang2026probguard,
  title = {{ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions}},
  author = {Xinzhe Huang and Biwu Yao and Kedong Xiu and Mengnan Zhao and Di Wang and Puning Zhao and Tianhang Zheng},
  year = {2026},
  month = aug,
  eprint = {2608.10621},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/f6c0d06532a9084259bdd4e94343c9a1e98e4f23}
}