September 2026Unreviewed
Audio Deepfake Detection Using Temporal Coherence Analysis
Justin D. Norman, Sarah Barrington
Abstract
The proliferation of AI-generated audio (so-called "deepfake" audio) poses significant threats to information integrity, from voice cloning fraud to synthetic music copyright disputes. We present a temporal coherence analysis framework built upon Contrastive Language-Audio Pretraining (CLAP) embeddings that spans speech, instrumental music, and music with vocals. By computing pairwise cosine similarities between audio segment embeddings and extracting statistical features from the resulting dist
Categories
Cite
@misc{norman2026audio,
title = {{Audio Deepfake Detection Using Temporal Coherence Analysis}},
author = {Justin D. Norman and Sarah Barrington},
year = {2026},
month = sep,
eprint = {2609.09489},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.09489}
}