Skip to content
Search
paperSeptember 2026Unreviewed

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

Vansh Wahi

Abstract

Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position is that it has not earned it. The judge should be demoted from oracle to advisor: its verdict becomes one input among several, and every change is gated instead by a deterministic verification layer the judge cannot override. We reached this posit

Categories

Cite

@misc{wahi2026llmasajudge,
  title = {{LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails}},
  author = {Vansh Wahi},
  year = {2026},
  month = sep,
  eprint = {2609.02246},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/d9bca7473eb20a165dd64c8710affd0e7263fdc6}
}