May 2026Unreviewed
Conformity Generates Collective Misalignment in AI Agents Societies
Giordano De Marzo, Alessandro Bellina, Claudio Castellano, Viola Priesemann, David Garcia
Abstract
Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here we show that populations of individually aligned AI agents can be driven into stable misaligned states through conformity dynamics. Simulating opinion dynamics across nine large language models and one hundred opinion pairs, we find that each agent's behavior
Categories
Cite
@misc{marzo2026conformity,
title = {{Conformity Generates Collective Misalignment in AI Agents Societies}},
author = {Giordano De Marzo and Alessandro Bellina and Claudio Castellano and Viola Priesemann and David Garcia},
year = {2026},
month = may,
eprint = {2605.10721},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.10721}
}