Skip to content
Search
paperSeptember 2026Unreviewed

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

Shenghan Zheng, Zonglin Di, Yimin Liu, Kyoung Whan Choe, Jiankai Sun, Heguang Lin, Penghao Jiang, Yifeng He, Xiao Cheng, Jicheng Wang, Wenbo Chen, Alex Yates, Yinzhe Zhao, Bingran You, Yuan Gao, Ayush Munot, Shubham Gaur, Zhe Ye, Hao Wang, Xiangyi Li, Dawn Song, Christophe Hauser

Abstract

LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task. Existing defenses rely largely on task-specific patches, prompt instructions, or post-hoc detectors. They do not

Categories

Cite

@misc{zheng2026benchshield,
  title = {{BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure}},
  author = {Shenghan Zheng and Zonglin Di and Yimin Liu and Kyoung Whan Choe and Jiankai Sun and Heguang Lin and Penghao Jiang and Yifeng He and Xiao Cheng and Jicheng Wang and Wenbo Chen and Alex Yates and Yinzhe Zhao and Bingran You and Yuan Gao and Ayush Munot and Shubham Gaur and Zhe Ye and Hao Wang and Xiangyi Li and Dawn Song and Christophe Hauser},
  year = {2026},
  month = sep,
  eprint = {2609.11028},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.11028}
}