← Back to search
paper llmsec-2026-00123
Self-Mined Hardness for Safety Fine-Tuning
Prakhar Gupta, Garv Shah, Donghua Zhang
2026-05
Abstract
Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's difficulty by how often the target model's own rollouts are judged harmful, then fine-tune on the hardest prompts paired with the model's own non-jailbroken rollouts. On Llama-3-8B-Instruct and Llama-3.2-3B-Instruct, this approach cuts the WildJailbreak attack success rate from 11.5% and 20.1% down to 1-3%, but pushes refusal on jailbreak-shaped beni
Categories
Cite This Resource
@article{llmsec202600123,
title = {Self-Mined Hardness for Safety Fine-Tuning},
author = {Prakhar Gupta and Garv Shah and Donghua Zhang},
year = {2026},
url = {https://arxiv.org/abs/2605.03226},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.03226