September 2026Unreviewed
ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
Hanling Wang, Chenlong Wei, Ling Xu, Hanyan Niu, Qi Cao, Shizhou Huang, Yang Yang, Xiaohui Zhu, Yao Zhu
Abstract
As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk methods that rely on surface tokens or output logits may suffer from weak initial signals, and internals-based detectors using dense representations may retain highly entangled and redundant safety-irrelevant information. It therefore remains unclear whether the earlies
Categories
Cite
@misc{wang2026foresight,
title = {{ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation}},
author = {Hanling Wang and Chenlong Wei and Ling Xu and Hanyan Niu and Qi Cao and Shizhou Huang and Yang Yang and Xiaohui Zhu and Yao Zhu},
year = {2026},
month = sep,
eprint = {2609.13737},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.13737}
}