June 2026Unreviewed
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
Marina Mancoridis, Zoë Hitzig
Abstract
Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pipelines depends on an implicit assumption: that the model applies relevant concepts the same way when it generates an output and later evaluates that output. We propose a new measure, generator-evaluator self-consistency, to test this assumption directly and apply it to 10 frontier models across 491 concepts. We find, f
Categories
Cite
@misc{mancoridis2026consistency,
title = {{The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes}},
author = {Marina Mancoridis and Zoë Hitzig},
year = {2026},
month = jun,
eprint = {2606.30653},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.30653}
}