Skip to content
Search
paperAugust 2026Unreviewed

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

Zhijie Zheng, Yu Li, Chen Qian, Yu-Qian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu

Abstract

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we i

Categories

Cite

@misc{zheng2026stepguard,
  title = {{StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing}},
  author = {Zhijie Zheng and Yu Li and Chen Qian and Yu-Qian Fu and Yanwei Fu and Lu Sheng and Jing Shao and Dongrui Liu},
  year = {2026},
  month = aug,
  eprint = {2608.24777},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/01d8aca918b46e65aaa866d65f619b80302510c6}
}