Skip to content
Search
paperSeptember 2026Unreviewed

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

Alejo López-Ávila, Iker García-Ferrero, Jezabel Garcia, Antonio Tiene, Román Orús

Abstract

Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-di

Categories

Cite

@misc{lopezavila2026safety,
  title = {{Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal}},
  author = {Alejo López-Ávila and Iker García-Ferrero and Jezabel Garcia and Antonio Tiene and Román Orús},
  year = {2026},
  month = sep,
  eprint = {2609.04482},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.04482}
}