Skip to content
Search
paperApril 2026Unreviewed

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs

Jaechul Roh, Amir Houmansadr

Abstract

Prior work shows that fine-tuning aligned models on benign data degrades safety in text and vision modalities, and that proximity to harmful content in representation space predicts which samples cause the most damage. However, existing analyses operate within a single, undifferentiated embedding space -- leaving open whether distinct input properties drive the vulnerability differently. Audio introduces a structurally richer problem: a benign sample can neighbor harmful content not only through

Categories

Cite

@misc{roh2026benign,
  title = {{Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs}},
  author = {Jaechul Roh and Amir Houmansadr},
  year = {2026},
  month = apr,
  eprint = {2604.16659},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.16659}
}