September 2026Unreviewed
Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning
Miso Kim, Georu Lee, Seungwon Jeong, Woojin Lee
Abstract
Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is inaccessible. We term this gap forget-set misalignment and identify two cases. In Under Unlearning, the forget set omits memorized information and leakage persists. In Out-of-Knowledge Unlearning, the algorithm is driven to "forget" knowledge the model never learned, pertu
Categories
Cite
@misc{kim2026confess,
title = {{Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning}},
author = {Miso Kim and Georu Lee and Seungwon Jeong and Woojin Lee},
year = {2026},
month = sep,
eprint = {2609.00605},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.00605}
}