Skip to content
Search
paperJune 2026Unreviewed

SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills

Ismail Hossain, Sai Puppala, Md Jahangir Alam, Tanzim Ahad, Sajedul Talukder

Abstract

Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent capabilities - remains largely unvetted. The gap we fill: existing scanners operate at the code layer and are structurally blind to instruction-layer and multi-agent risk - natural-language directives that hijack an agent, exfiltrate data through encoded side channels, or chain harm across pipelines - so what is needed is a semantic, multi-dimensiona

Categories

Cite

@misc{hossain2026skillvetbench,
  title = {{SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills}},
  author = {Ismail Hossain and Sai Puppala and Md Jahangir Alam and Tanzim Ahad and Sajedul Talukder},
  year = {2026},
  month = jun,
  eprint = {2606.15899},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.15899}
}