← Back to search
paper llmsec-2026-00101
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
Zhenhao Xu, Wenhan Chang, Yichuan Chen, Yuxin Fang, Junhao Liu, Tianqing Zhu
2026-05
Abstract
Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify model weights and must instead intervene at inference time. This setting creates three practical challenges: harmful intent may be hidden by educational or role-play framing, deep safety analysis can introduce non-trivial latency, and long adversarial contexts can dilute the local cues that simpler filters rely on. These c
Categories
Cite This Resource
@article{llmsec202600101,
title = {Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis},
author = {Zhenhao Xu and Wenhan Chang and Yichuan Chen and Yuxin Fang and Junhao Liu and Tianqing Zhu},
year = {2026},
url = {https://arxiv.org/abs/2605.11664},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.11664