Skip to content
Search
paperJune 2026Unreviewed

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents

Yuchuan Tian, Mengyu Zheng, Haocheng Mei, Ye Yuan, Chao Xu, Xinghao Chen, Hanting Chen, Yu Wang

Abstract

Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects. Existing evaluations often collapse these stages into a single attack success rate, making it difficult to tell whether a model merely agreed with an attacker or actually produced observable harm. We introduce SafeClawBench, a staged benchmark for tool-using agent security

Categories

Cite

@misc{tian2026safeclawbench,
  title = {{SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents}},
  author = {Yuchuan Tian and Mengyu Zheng and Haocheng Mei and Ye Yuan and Chao Xu and Xinghao Chen and Hanting Chen and Yu Wang},
  year = {2026},
  month = jun,
  eprint = {2606.18356},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.18356}
}