← Back to search
paper llmsec-2026-00123

Self-Mined Hardness for Safety Fine-Tuning

Prakhar Gupta, Garv Shah, Donghua Zhang

2026-05

Abstract

Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's difficulty by how often the target model's own rollouts are judged harmful, then fine-tune on the hardest prompts paired with the model's own non-jailbroken rollouts. On Llama-3-8B-Instruct and Llama-3.2-3B-Instruct, this approach cuts the WildJailbreak attack success rate from 11.5% and 20.1% down to 1-3%, but pushes refusal on jailbreak-shaped beni

Categories

Cite This Resource

@article{llmsec202600123,
  title = {Self-Mined Hardness for Safety Fine-Tuning},
  author = {Prakhar Gupta and Garv Shah and Donghua Zhang},
  year = {2026},
  url = {https://arxiv.org/abs/2605.03226},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.03226