September 2026Unreviewed
Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models
Kunlin Cai, Kaiyuan Zhang, Zihang Xiang, Jinghuai Zhang, Abeer Alwan, Fnu Suya, Yuan Tian
Abstract
Text-to-Speech (TTS) foundation models are increasingly fine-tuned on private datasets to synthesize highly personalized voices, introducing severe privacy risks by exposing both biometric identities and sensitive speech content. Existing black-box membership inference attacks (MIAs) follow a two-stage pipeline of query generation and representation engineering, both of which face unique challenges when adapted to TTS. For query generation, dual conditioning on synthesis text and reference speec
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM02Sensitive Information Disclosure
MITRE ATLAS
- AML.T0024.000Infer Training Data Membership
Suggested from the entry's categories.
Cite
@misc{cai2026hearing,
title = {{Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models}},
author = {Kunlin Cai and Kaiyuan Zhang and Zihang Xiang and Jinghuai Zhang and Abeer Alwan and Fnu Suya and Yuan Tian},
year = {2026},
month = sep,
eprint = {2609.01723},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.01723}
}