June 2026Unreviewed
An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic
Shuze Liu, Qianwen Guo, Yushun Dong
Abstract
Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service security. However, individual extraction queries often resemble benign requests, and existing evaluations often focus on single-query anomaly scoring or pure benign-versus-attacker user settings. We formulate model extraction monitoring as benign-calibrated traffic-window distribution testing and show that an embarrassingly simple detector is effec
Categories
Framework mappings
MITRE ATLAS
- AML.T0024.002Extract AI Model
Suggested from the entry's categories.
Cite
@misc{liu2026embarrassingly,
title = {{An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic}},
author = {Shuze Liu and Qianwen Guo and Yushun Dong},
year = {2026},
month = jun,
eprint = {2606.05725},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.05725}
}