June 2026Unreviewed
SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills
Ismail Hossain, Sai Puppala, Md Jahangir Alam, Tanzim Ahad, Sajedul Talukder
Abstract
Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent capabilities - remains largely unvetted. The gap we fill: existing scanners operate at the code layer and are structurally blind to instruction-layer and multi-agent risk - natural-language directives that hijack an agent, exfiltrate data through encoded side channels, or chain harm across pipelines - so what is needed is a semantic, multi-dimensiona
Categories
Cite
@misc{hossain2026skillvetbench,
title = {{SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills}},
author = {Ismail Hossain and Sai Puppala and Md Jahangir Alam and Tanzim Ahad and Sajedul Talukder},
year = {2026},
month = jun,
eprint = {2606.15899},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.15899}
}