May 2026Unreviewed
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
Kevin Kuo, Chhavi Yadav, Virginia Smith
Abstract
Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this pape
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{kuo2026openweight,
title = {{Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks}},
author = {Kevin Kuo and Chhavi Yadav and Virginia Smith},
year = {2026},
month = may,
eprint = {2605.26526},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.26526}
}