June 2026Unreviewed
MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills
Wenbo Guo, Wei Zeng, Chengwei Liu, Xiaojun Jia, Yijia Xu, Lei Tang, Yong Fang, Yang Liu
Abstract
AI coding agents such as Claude Code and Gemini CLI increasingly extend themselves with third-party skills: markdown packages bundling natural-language instructions, executable scripts, and tool permissions. Because a skill is at once code and agent-facing instruction, it introduces a supply chain dependency whose risk is neither pure code nor pure prompt. Detection tools have never been measured against verified ground truth spanning this hybrid space, leaving their effectiveness unknown and wi
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM03Supply Chain
MITRE ATLAS
- AML.T0010AI Supply Chain Compromise
Suggested from the entry's categories.
Cite
@misc{guo2026malskillbench,
title = {{MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills}},
author = {Wenbo Guo and Wei Zeng and Chengwei Liu and Xiaojun Jia and Yijia Xu and Lei Tang and Yong Fang and Yang Liu},
year = {2026},
month = jun,
eprint = {2606.07131},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.07131}
}