← Back to search
paper llmsec-2026-00171

Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs

Jingshen Zhang, Bo Wang, Yanlin Fu, Dongming Zhao, Ruifang He, Yuexian Hou, Zifei Yu

2026-05

Abstract

In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily triggered by explicit input-level stimuli, self-debiasing mechanisms can involve generation-time intrinsic correction that are not directly reducible to surface-level prompt. Motivated by conflict-monitoring and response-inhibition accounts in cognitive neuroscience, we propose COCO, a contrastive causal method design

Cite This Resource

@article{llmsec202600171,
  title = {Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs},
  author = {Jingshen Zhang and Bo Wang and Yanlin Fu and Dongming Zhao and Ruifang He and Yuexian Hou and Zifei Yu},
  year = {2026},
  url = {https://arxiv.org/abs/2605.09647},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.09647