August 2026Unreviewed
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
Zhijie Zheng, Yu Li, Chen Qian, Yu-Qian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu
Abstract
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we i
Categories
Cite
@misc{zheng2026stepguard,
title = {{StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing}},
author = {Zhijie Zheng and Yu Li and Chen Qian and Yu-Qian Fu and Yanwei Fu and Lu Sheng and Jing Shao and Dongrui Liu},
year = {2026},
month = aug,
eprint = {2608.24777},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/01d8aca918b46e65aaa866d65f619b80302510c6}
}