May 2026Unreviewed
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
Zhenhao Xu, Wenhan Chang, Yichuan Chen, Yuxin Fang, Junhao Liu, Tianqing Zhu
Abstract
Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify model weights and must instead intervene at inference time. This setting creates three practical challenges: harmful intent may be hidden by educational or role-play framing, deep safety analysis can introduce non-trivial latency, and long adversarial contexts can dilute the local cues that simpler filters rely on. These c
Categories
Cite
@misc{xu2026safety,
title = {{Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis}},
author = {Zhenhao Xu and Wenhan Chang and Yichuan Chen and Yuxin Fang and Junhao Liu and Tianqing Zhu},
year = {2026},
month = may,
eprint = {2605.11664},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.11664}
}