Skip to content
Search
paperApril 2026Unreviewed

Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries

Ki Sen Hung, Xi Yang, Chang Liu, Haoran Li, Kejiang Chen, Changxuan Fan, Tsun On Kwok, Weiming Zhang, Xiaomeng Li, Yangqiu Song

Abstract

A central goal of LLM alignment is to balance helpfulness with harmlessness, yet these objectives conflict when the same knowledge serves both legitimate and malicious purposes. This tension is amplified by context-sensitive alignment: we observe that domain-specific contexts (e.g., chemistry) selectively relax defenses for domain-relevant harmful knowledge, while safety-research contexts (e.g., jailbreak studies) trigger broader relaxation spanning all harm categories. To systematically exploit

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hung2026into,
  title = {{Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries}},
  author = {Ki Sen Hung and Xi Yang and Chang Liu and Haoran Li and Kejiang Chen and Changxuan Fan and Tsun On Kwok and Weiming Zhang and Xiaomeng Li and Yangqiu Song},
  year = {2026},
  month = apr,
  eprint = {2604.15717},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.15717}
}