August 2026Unreviewed
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
Víctor Gallego
Abstract
Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU-kernel-optimization suites with held-out generalization gates: Metal-Sci (10 scientific-compute tasks) and Metal-ZK (12 zero-knowledge/cryptographic tasks), in which three frontier LLMs (Opus 4.7, Gemini 3.1 Pro, GPT-5.5) propose Metal kernels inside a $(1{+}1)$ evolutionary loop with rich feedback. Although no model is prompted to act a
Categories
Cite
@misc{gallego2026gaming,
title = {{Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure}},
author = {Víctor Gallego},
year = {2026},
month = aug,
eprint = {2608.08722},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.08722}
}