May 2026Unreviewed
No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
Ying Li, Hongbo Wen, Yanju Chen, Hanzhi Liu, Yuan Tian, Yu Feng
Abstract
LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, but because the skill it invoked broke its own declared safety rules. We call these specification violations: benign inputs cause a skill to breach the natural-language guardrails in its own specification, typically because the guardrail's semantics are undefined for autonomous execution, or because the implementation silently ignores the documented
Categories
Cite
@misc{li2026no,
title = {{No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills}},
author = {Ying Li and Hongbo Wen and Yanju Chen and Hanzhi Liu and Yuan Tian and Yu Feng},
year = {2026},
month = may,
eprint = {2605.13044},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.13044}
}