April 2026Unreviewed
Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries
Ki Sen Hung, Xi Yang, Chang Liu, Haoran Li, Kejiang Chen, Changxuan Fan, Tsun On Kwok, Weiming Zhang, Xiaomeng Li, Yangqiu Song
Abstract
A central goal of LLM alignment is to balance helpfulness with harmlessness, yet these objectives conflict when the same knowledge serves both legitimate and malicious purposes. This tension is amplified by context-sensitive alignment: we observe that domain-specific contexts (e.g., chemistry) selectively relax defenses for domain-relevant harmful knowledge, while safety-research contexts (e.g., jailbreak studies) trigger broader relaxation spanning all harm categories. To systematically exploit
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{hung2026into,
title = {{Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries}},
author = {Ki Sen Hung and Xi Yang and Chang Liu and Haoran Li and Kejiang Chen and Changxuan Fan and Tsun On Kwok and Weiming Zhang and Xiaomeng Li and Yangqiu Song},
year = {2026},
month = apr,
eprint = {2604.15717},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.15717}
}