Search papers, labs, and topics across Lattice.
To tackle synthetic audio detection across speech and music, this work evaluates temporal coherence by extracting statistical properties from pairwise cosine similarity distributions over segment-level CLAP embeddings. Lightweight ensemble classifiers trained on these metric distributions achieve competitive performance with deep networks while offering full interpretability. Crucially, the analysis reveals severe domain-shift pitfalls: 21 of 29 statistical features invert their discriminative direction in-the-wild, and entropy discriminates in opposite directions when evaluating synthetic speech versus synthetic music.
In-the-wild audio deepfakes break standard forensic assumptions: 21 of 29 temporal coherence metrics invert their discriminative direction outside training distributions, with markers like entropy flipping signs entirely between synthetic speech and music.
The proliferation of AI-generated audio (so-called"deepfake"audio) poses significant threats to information integrity, from voice cloning fraud to synthetic music copyright disputes. We present a temporal coherence analysis framework built upon Contrastive Language-Audio Pretraining (CLAP) embeddings that spans speech, instrumental music, and music with vocals. By computing pairwise cosine similarities between audio segment embeddings and extracting statistical features from the resulting distributions, we train lightweight ensemble classifiers that reliably distinguish authentic from synthetic audio. Our work provides an interpretable, computationally efficient alternative to common deep learning methods while still achieving competitive performance across speech and music domains. Further, we reveal two notable empirical findings about audio deepfakes: (1) a feature-label inversion phenomenon in which 21 of 29 statistical features reverse their discriminative direction between training and in-the-wild deployment, and (2) a speech--music direction reversal in which entropy discriminates in opposite directions for speech and music deepfakes.