June 2026Unreviewed
A Unified Evaluation Framework for Utility and Privacy Risks of LLM-Generated Synthetic Text Data
Lubana Isaoglu, Zeynep Orman
Abstract
The increasing use of Large Language Models (LLMs) has enabled the generation of high-quality synthetic text, providing a potential alternative to sensitive real-world datasets in domains where privacy concerns limit data sharing. However, synthetic text is not inherently privacy safe. Fine-tuning generative models on domain-specific data can enhance semantic fidelity while simultaneously increasing the risk of memorization and information leaks. In this work, we propose a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0024.000Infer Training Data Membership
Suggested from the entry's categories.
Cite
@misc{isaoglu2026unified,
title = {{A Unified Evaluation Framework for Utility and Privacy Risks of LLM-Generated Synthetic Text Data}},
author = {Lubana Isaoglu and Zeynep Orman},
year = {2026},
month = jun,
doi = {10.64808/engineeringperspective.1910777},
url = {https://doi.org/10.64808/engineeringperspective.1910777}
}