Skip to content
Search
paperMay 2026Unreviewed

Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs

Jingshen Zhang, Bo Wang, Yanlin Fu, Dongming Zhao, Ruifang He, Yuexian Hou, Zifei Yu

Abstract

In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily triggered by explicit input-level stimuli, self-debiasing mechanisms can involve generation-time intrinsic correction that are not directly reducible to surface-level prompt. Motivated by conflict-monitoring and response-inhibition accounts in cognitive neuroscience, we propose COCO, a contrastive causal method design

Categories

Cite

@misc{zhang2026modeling,
  title = {{Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs}},
  author = {Jingshen Zhang and Bo Wang and Yanlin Fu and Dongming Zhao and Ruifang He and Yuexian Hou and Zifei Yu},
  year = {2026},
  month = may,
  eprint = {2605.09647},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.09647}
}