Skip to content
Search
paperSeptember 2026Unreviewed

Audio Deepfake Detection Using Temporal Coherence Analysis

Justin D. Norman, Sarah Barrington

Abstract

The proliferation of AI-generated audio (so-called "deepfake" audio) poses significant threats to information integrity, from voice cloning fraud to synthetic music copyright disputes. We present a temporal coherence analysis framework built upon Contrastive Language-Audio Pretraining (CLAP) embeddings that spans speech, instrumental music, and music with vocals. By computing pairwise cosine similarities between audio segment embeddings and extracting statistical features from the resulting dist

Categories

Cite

@misc{norman2026audio,
  title = {{Audio Deepfake Detection Using Temporal Coherence Analysis}},
  author = {Justin D. Norman and Sarah Barrington},
  year = {2026},
  month = sep,
  eprint = {2609.09489},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.09489}
}