August 2026Unreviewed
A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families
Yu Zhe, Yiting Tan, Jun-Hao Wei, Wan Chen
Abstract
Model merging enables combining multiple fine-tuned models without additional training, but its safety implications remain poorly understood. Prior work primarily attributes merging risks to unsafe constituent models, implicitly assuming that merging individually aligned models preserves safety. In contrast, we show that model merging reveals a previously overlooked jailbreak risk rooted in the pretrained foundation model, even when all constituent models are individually safety-aligned. Motivat
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhe2026single,
title = {{A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families}},
author = {Yu Zhe and Yiting Tan and Jun-Hao Wei and Wan Chen},
year = {2026},
month = aug,
eprint = {2608.26506},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/0f2d5b2bfa30e140a8aa2849861f3bcbca2bd73e}
}