August 2026Unreviewed
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
Mark Russinovich
Abstract
Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified.
Categories
Cite
@misc{russinovich2026fools,
title = {{Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models}},
author = {Mark Russinovich},
year = {2026},
month = aug,
eprint = {2608.17202},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.17202}
}