September 2026Unreviewed
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
Vansh Wahi
Abstract
Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position is that it has not earned it. The judge should be demoted from oracle to advisor: its verdict becomes one input among several, and every change is gated instead by a deterministic verification layer the judge cannot override. We reached this posit
Categories
Cite
@misc{wahi2026llmasajudge,
title = {{LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails}},
author = {Vansh Wahi},
year = {2026},
month = sep,
eprint = {2609.02246},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/d9bca7473eb20a165dd64c8710affd0e7263fdc6}
}