September 2026Unreviewed
MOLE: Detecting Insider Threats in AI Agents
Aashiq Muhamed, Virginia Smith
Abstract
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tok
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{muhamed2026mole,
title = {{MOLE: Detecting Insider Threats in AI Agents}},
author = {Aashiq Muhamed and Virginia Smith},
year = {2026},
month = sep,
eprint = {2609.06966},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.06966}
}