← Back to search
paper llmsec-2026-00171
Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs
Jingshen Zhang, Bo Wang, Yanlin Fu, Dongming Zhao, Ruifang He, Yuexian Hou, Zifei Yu
2026-05
Abstract
In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily triggered by explicit input-level stimuli, self-debiasing mechanisms can involve generation-time intrinsic correction that are not directly reducible to surface-level prompt. Motivated by conflict-monitoring and response-inhibition accounts in cognitive neuroscience, we propose COCO, a contrastive causal method design
Categories
Cite This Resource
@article{llmsec202600171,
title = {Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs},
author = {Jingshen Zhang and Bo Wang and Yanlin Fu and Dongming Zhao and Ruifang He and Yuexian Hou and Zifei Yu},
year = {2026},
url = {https://arxiv.org/abs/2605.09647},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.09647