Skip to content
Search
paperSeptember 2026Unreviewed

On the Impact of Anonymization on the Performance of Large Language Models

Tobias Deußer, Max Hahnbück, Lorenz Sparrenberg, Tobias Uelwer, Christian Bauckhage, Rafet Sifa

Abstract

As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper presents a systematic empirical study of the trade-off between privacy and performance. We evaluate five prominent language models across eleven diverse benchmarks, comparing their performance on original versus pseudonymized input

Categories

Cite

@misc{deuer2026impact,
  title = {{On the Impact of Anonymization on the Performance of Large Language Models}},
  author = {Tobias Deußer and Max Hahnbück and Lorenz Sparrenberg and Tobias Uelwer and Christian Bauckhage and Rafet Sifa},
  year = {2026},
  month = sep,
  eprint = {2609.11335},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.11335}
}