Skip to content
Search
paperMay 2026Unreviewed

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

Kevin Kuo, Chhavi Yadav, Virginia Smith

Abstract

Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this pape

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{kuo2026openweight,
  title = {{Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks}},
  author = {Kevin Kuo and Chhavi Yadav and Virginia Smith},
  year = {2026},
  month = may,
  eprint = {2605.26526},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.26526}
}