Skip to content
Search
paperSeptember 2026Unreviewed

MOLE: Detecting Insider Threats in AI Agents

Aashiq Muhamed, Virginia Smith

Abstract

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tok

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{muhamed2026mole,
  title = {{MOLE: Detecting Insider Threats in AI Agents}},
  author = {Aashiq Muhamed and Virginia Smith},
  year = {2026},
  month = sep,
  eprint = {2609.06966},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.06966}
}