May 2026Unreviewed
Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs
Jingshen Zhang, Bo Wang, Yanlin Fu, Dongming Zhao, Ruifang He, Yuexian Hou, Zifei Yu
Abstract
In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily triggered by explicit input-level stimuli, self-debiasing mechanisms can involve generation-time intrinsic correction that are not directly reducible to surface-level prompt. Motivated by conflict-monitoring and response-inhibition accounts in cognitive neuroscience, we propose COCO, a contrastive causal method design
Categories
Cite
@misc{zhang2026modeling,
title = {{Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs}},
author = {Jingshen Zhang and Bo Wang and Yanlin Fu and Dongming Zhao and Ruifang He and Yuexian Hou and Zifei Yu},
year = {2026},
month = may,
eprint = {2605.09647},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.09647}
}