Skip to content
Search
paperAugust 2026Unreviewed

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

Yu Zhe, Yiting Tan, Jun-Hao Wei, Wan Chen

Abstract

Model merging enables combining multiple fine-tuned models without additional training, but its safety implications remain poorly understood. Prior work primarily attributes merging risks to unsafe constituent models, implicitly assuming that merging individually aligned models preserves safety. In contrast, we show that model merging reveals a previously overlooked jailbreak risk rooted in the pretrained foundation model, even when all constituent models are individually safety-aligned. Motivat

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhe2026single,
  title = {{A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families}},
  author = {Yu Zhe and Yiting Tan and Jun-Hao Wei and Wan Chen},
  year = {2026},
  month = aug,
  eprint = {2608.26506},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/0f2d5b2bfa30e140a8aa2849861f3bcbca2bd73e}
}