Skip to content
Search
paperJune 2026Unreviewed

MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills

Wenbo Guo, Wei Zeng, Chengwei Liu, Xiaojun Jia, Yijia Xu, Lei Tang, Yong Fang, Yang Liu

Abstract

AI coding agents such as Claude Code and Gemini CLI increasingly extend themselves with third-party skills: markdown packages bundling natural-language instructions, executable scripts, and tool permissions. Because a skill is at once code and agent-facing instruction, it introduces a supply chain dependency whose risk is neither pure code nor pure prompt. Detection tools have never been measured against verified ground truth spanning this hybrid space, leaving their effectiveness unknown and wi

Categories

Framework mappings

MITRE ATLAS
  • AML.T0010AI Supply Chain Compromise

Suggested from the entry's categories.

Cite

@misc{guo2026malskillbench,
  title = {{MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills}},
  author = {Wenbo Guo and Wei Zeng and Chengwei Liu and Xiaojun Jia and Yijia Xu and Lei Tang and Yong Fang and Yang Liu},
  year = {2026},
  month = jun,
  eprint = {2606.07131},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.07131}
}