August 2026Unreviewed
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou
Abstract
Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful prompts scored for compliance. We argue this is misleading. Because safety alignment is"shallow,"concentrated in the first few ge
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{ma2026when,
title = {{When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs}},
author = {Yu Ma and Hongli Shi and Jing Li and Xinran Xu and Weiwei Hou},
year = {2026},
month = aug,
eprint = {2608.08542},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/4ef3b426aaf844b65a8b487baa1a87b1aec4c55e}
}