Skip to content
Search
paperSeptember 2026Unreviewed

RISA: Response Inspection and Selective Actions for Refusal Calibration in Large Language Models

Wenhan Chang, Tianqing Zhu, Ping Xiong, Shiyi Liao, Wanlei Zhou

Abstract

Reliable refusal behavior requires Large Language Models (LLMs) to reject harmful prompts with only answering benign ones. Incorrect refusal behavior can either expose users to harmful responses or prevent users from obtaining useful answers. Training-time alignment improves refusal behavior by updating model parameters with safety data, but requires additional computation and training. In contrast, inference-time alignment aims to modify LLM behavior during inference without updating the underl

Categories

Cite

@misc{chang2026risa,
  title = {{RISA: Response Inspection and Selective Actions for Refusal Calibration in Large Language Models}},
  author = {Wenhan Chang and Tianqing Zhu and Ping Xiong and Shiyi Liao and Wanlei Zhou},
  year = {2026},
  month = sep,
  eprint = {2609.00790},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.00790}
}