Skip to content
Search
paperJune 2026Unreviewed

The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes

Marina Mancoridis, Zoë Hitzig

Abstract

Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pipelines depends on an implicit assumption: that the model applies relevant concepts the same way when it generates an output and later evaluates that output. We propose a new measure, generator-evaluator self-consistency, to test this assumption directly and apply it to 10 frontier models across 491 concepts. We find, f

Categories

Cite

@misc{mancoridis2026consistency,
  title = {{The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes}},
  author = {Marina Mancoridis and Zoë Hitzig},
  year = {2026},
  month = jun,
  eprint = {2606.30653},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.30653}
}