Skip to content
Search
paperApril 2026Unreviewed

Owner-Harm: A Missing Threat Model for AI Agent Safety

Dongcheng Zhang, Yiqing Jiang

Abstract

Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a distinct and commercially consequential threat category: agents harming their own deployers. Real-world incidents illustrate the gap: Slack AI credential exfiltration (Aug 2024), Microsoft 365 Copilot calendar-injection leaks (Jan 2024), and a Meta agent unauthorized forum post exposing operational data (Mar 2026). We propose Owner-Harm, a formal th

Categories

Cite

@misc{zhang2026ownerharm,
  title = {{Owner-Harm: A Missing Threat Model for AI Agent Safety}},
  author = {Dongcheng Zhang and Yiqing Jiang},
  year = {2026},
  month = apr,
  eprint = {2604.18658},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.18658}
}