Skip to content
Search
paperAugust 2026Unreviewed

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

Chen Chen, Yao-Lin Chen, Xue-Han Sun, Juan Lin, Xueluan Gong, Yu-Heng Zheng, Qian Wang, Kwok-Yan Lam

Abstract

Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not specifically target LLM judges and provide limited support for multiple evaluation protocols under restricted query budgets. In this study, we propose JUDGESTEALER, the first query-efficient model extraction framework for r

Categories

Framework mappings

MITRE ATLAS
  • AML.T0024.002Extract AI Model

Suggested from the entry's categories.

Cite

@misc{chen2026judgestealer,
  title = {{JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols}},
  author = {Chen Chen and Yao-Lin Chen and Xue-Han Sun and Juan Lin and Xueluan Gong and Yu-Heng Zheng and Qian Wang and Kwok-Yan Lam},
  year = {2026},
  month = aug,
  eprint = {2608.26982},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/0becf5302b4eabed685d226ee6ffbf1f3a06764e}
}