September 2026Unreviewed
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
Alejo López-Ávila, Iker García-Ferrero, Jezabel Garcia, Antonio Tiene, Román Orús
Abstract
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-di
Categories
Cite
@misc{lopezavila2026safety,
title = {{Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal}},
author = {Alejo López-Ávila and Iker García-Ferrero and Jezabel Garcia and Antonio Tiene and Román Orús},
year = {2026},
month = sep,
eprint = {2609.04482},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.04482}
}