September 2026Unreviewed
On the Impact of Anonymization on the Performance of Large Language Models
Tobias Deußer, Max Hahnbück, Lorenz Sparrenberg, Tobias Uelwer, Christian Bauckhage, Rafet Sifa
Abstract
As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper presents a systematic empirical study of the trade-off between privacy and performance. We evaluate five prominent language models across eleven diverse benchmarks, comparing their performance on original versus pseudonymized input
Categories
Cite
@misc{deuer2026impact,
title = {{On the Impact of Anonymization on the Performance of Large Language Models}},
author = {Tobias Deußer and Max Hahnbück and Lorenz Sparrenberg and Tobias Uelwer and Christian Bauckhage and Rafet Sifa},
year = {2026},
month = sep,
eprint = {2609.11335},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.11335}
}