← Back to search
paper llmsec-2026-00101

Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis

Zhenhao Xu, Wenhan Chang, Yichuan Chen, Yuxin Fang, Junhao Liu, Tianqing Zhu

2026-05

Abstract

Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify model weights and must instead intervene at inference time. This setting creates three practical challenges: harmful intent may be hidden by educational or role-play framing, deep safety analysis can introduce non-trivial latency, and long adversarial contexts can dilute the local cues that simpler filters rely on. These c

Cite This Resource

@article{llmsec202600101,
  title = {Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis},
  author = {Zhenhao Xu and Wenhan Chang and Yichuan Chen and Yuxin Fang and Junhao Liu and Tianqing Zhu},
  year = {2026},
  url = {https://arxiv.org/abs/2605.11664},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.11664