Skip to content
Search
paperSeptember 2026Unreviewed

What a Model Refuses, a State Fears: How Authoritarian Information Control Reproduces in Language-Model Guardrails

Meng-Lin Liu, Yao Yu, Tong Wu, Chun-Ran Zhang, Ge Shi

Abstract

As large language models become the front door to political information, what they refuse to discuss becomes a new instrument of information control. We argue that a model's guardrail encodes not a universal notion of harm but the political threat model of the state that governs its developer, and we derive the expected structure of that control from the comparative study of how authoritarian regimes censor. Across ten models and three languages, Chinese guardrails carry its signatures: they ans

Categories

Cite

@misc{liu2026what,
  title = {{What a Model Refuses, a State Fears: How Authoritarian Information Control Reproduces in Language-Model Guardrails}},
  author = {Meng-Lin Liu and Yao Yu and Tong Wu and Chun-Ran Zhang and Ge Shi},
  year = {2026},
  month = sep,
  eprint = {2609.07507},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/475fe2ba2121de45c572735d75d9efa667deaa30}
}