Skip to content
Search
paperAugust 2026Unreviewed

Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

Mark Russinovich

Abstract

Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified.

Categories

Cite

@misc{russinovich2026fools,
  title = {{Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models}},
  author = {Mark Russinovich},
  year = {2026},
  month = aug,
  eprint = {2608.17202},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.17202}
}