August 2026Unreviewed
JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols
Chen Chen, Yao-Lin Chen, Xue-Han Sun, Juan Lin, Xueluan Gong, Yu-Heng Zheng, Qian Wang, Kwok-Yan Lam
Abstract
Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not specifically target LLM judges and provide limited support for multiple evaluation protocols under restricted query budgets. In this study, we propose JUDGESTEALER, the first query-efficient model extraction framework for r
Categories
Framework mappings
MITRE ATLAS
- AML.T0024.002Extract AI Model
Suggested from the entry's categories.
Cite
@misc{chen2026judgestealer,
title = {{JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols}},
author = {Chen Chen and Yao-Lin Chen and Xue-Han Sun and Juan Lin and Xueluan Gong and Yu-Heng Zheng and Qian Wang and Kwok-Yan Lam},
year = {2026},
month = aug,
eprint = {2608.26982},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/0becf5302b4eabed685d226ee6ffbf1f3a06764e}
}