Skip to content
Search
paperMay 2026Unreviewed

Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis

Zhenhao Xu, Wenhan Chang, Yichuan Chen, Yuxin Fang, Junhao Liu, Tianqing Zhu

Abstract

Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify model weights and must instead intervene at inference time. This setting creates three practical challenges: harmful intent may be hidden by educational or role-play framing, deep safety analysis can introduce non-trivial latency, and long adversarial contexts can dilute the local cues that simpler filters rely on. These c

Categories

Cite

@misc{xu2026safety,
  title = {{Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis}},
  author = {Zhenhao Xu and Wenhan Chang and Yichuan Chen and Yuxin Fang and Junhao Liu and Tianqing Zhu},
  year = {2026},
  month = may,
  eprint = {2605.11664},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.11664}
}