Skip to content
Search
paperJune 2026Unreviewed

Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot

Yuyang Dai, Yushun Dong

Abstract

Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late or degrade utility for legitimate users. We propose \textbf{Knowledge Trap}, a defense that redirects extraction attacks toward low-transferability knowledge through a \emph{Honeypot Knowledge Graph} (HKG) and breadcrumb-guided exploration. Instead of blocking queries or perturbing outputs, Knowledge Trap consumes the attacker's limited query budget on knowled

Categories

Framework mappings

MITRE ATLAS
  • AML.T0024.002Extract AI Model

Suggested from the entry's categories.

Cite

@misc{dai2026let,
  title = {{Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot}},
  author = {Yuyang Dai and Yushun Dong},
  year = {2026},
  month = jun,
  eprint = {2606.15810},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.15810}
}