Skip to content
Search
paperJune 2026Unreviewed

A Unified Evaluation Framework for Utility and Privacy Risks of LLM-Generated Synthetic Text Data

Lubana Isaoglu, Zeynep Orman

Abstract

The increasing use of Large Language Models (LLMs) has enabled the generation of high-quality synthetic text, providing a potential alternative to sensitive real-world datasets in domains where privacy concerns limit data sharing. However, synthetic text is not inherently privacy safe. Fine-tuning generative models on domain-specific data can enhance semantic fidelity while simultaneously increasing the risk of memorization and information leaks. In this work, we propose a

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0024.000Infer Training Data Membership

Suggested from the entry's categories.

Cite

@misc{isaoglu2026unified,
  title = {{A Unified Evaluation Framework for Utility and Privacy Risks of LLM-Generated Synthetic Text Data}},
  author = {Lubana Isaoglu and Zeynep Orman},
  year = {2026},
  month = jun,
  doi = {10.64808/engineeringperspective.1910777},
  url = {https://doi.org/10.64808/engineeringperspective.1910777}
}