Skip to content
Search
paperAugust 2026Unreviewed

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou

Abstract

Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful prompts scored for compliance. We argue this is misleading. Because safety alignment is"shallow,"concentrated in the first few ge

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{ma2026when,
  title = {{When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs}},
  author = {Yu Ma and Hongli Shi and Jing Li and Xinran Xu and Weiwei Hou},
  year = {2026},
  month = aug,
  eprint = {2608.08542},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/4ef3b426aaf844b65a8b487baa1a87b1aec4c55e}
}