April 2026Unreviewed
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
Jaechul Roh, Amir Houmansadr
Abstract
Prior work shows that fine-tuning aligned models on benign data degrades safety in text and vision modalities, and that proximity to harmful content in representation space predicts which samples cause the most damage. However, existing analyses operate within a single, undifferentiated embedding space -- leaving open whether distinct input properties drive the vulnerability differently. Audio introduces a structurally richer problem: a benign sample can neighbor harmful content not only through
Categories
Cite
@misc{roh2026benign,
title = {{Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs}},
author = {Jaechul Roh and Amir Houmansadr},
year = {2026},
month = apr,
eprint = {2604.16659},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.16659}
}