September 2026Unreviewed
Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random
Cris Huynh
Abstract
Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using a forced-choice signalling task abstracted from the board game Deception: Murder in Hong Kong. In this environment, the reference points against which a coordinate should be judged (a fit-maximising strategy, a posterior-maximising strategy, and uniform random selecti
Categories
Cite
@misc{huynh2026consistency,
title = {{Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random}},
author = {Cris Huynh},
year = {2026},
month = sep,
eprint = {2609.00576},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.00576}
}