Skip to content
Search
paperSeptember 2026Unreviewed

Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models

Kunlin Cai, Kaiyuan Zhang, Zihang Xiang, Jinghuai Zhang, Abeer Alwan, Fnu Suya, Yuan Tian

Abstract

Text-to-Speech (TTS) foundation models are increasingly fine-tuned on private datasets to synthesize highly personalized voices, introducing severe privacy risks by exposing both biometric identities and sensitive speech content. Existing black-box membership inference attacks (MIAs) follow a two-stage pipeline of query generation and representation engineering, both of which face unique challenges when adapted to TTS. For query generation, dual conditioning on synthesis text and reference speec

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure
MITRE ATLAS
  • AML.T0024.000Infer Training Data Membership

Suggested from the entry's categories.

Cite

@misc{cai2026hearing,
  title = {{Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models}},
  author = {Kunlin Cai and Kaiyuan Zhang and Zihang Xiang and Jinghuai Zhang and Abeer Alwan and Fnu Suya and Yuan Tian},
  year = {2026},
  month = sep,
  eprint = {2609.01723},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.01723}
}