June 2026Unreviewed
SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents
Yuchuan Tian, Mengyu Zheng, Haocheng Mei, Ye Yuan, Chao Xu, Xinghao Chen, Hanting Chen, Yu Wang
Abstract
Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects. Existing evaluations often collapse these stages into a single attack success rate, making it difficult to tell whether a model merely agreed with an attacker or actually produced observable harm. We introduce SafeClawBench, a staged benchmark for tool-using agent security
Categories
Cite
@misc{tian2026safeclawbench,
title = {{SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents}},
author = {Yuchuan Tian and Mengyu Zheng and Haocheng Mei and Ye Yuan and Chao Xu and Xinghao Chen and Hanting Chen and Yu Wang},
year = {2026},
month = jun,
eprint = {2606.18356},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.18356}
}