Search papers, labs, and topics across Lattice.
This study evaluates five frozen transformer encoders鈥擜ST, Vaani-FastConformer, Wav2vec2, Whisper, and BEATs鈥攁cross 22 Indic languages to assess their effectiveness in distinguishing spontaneous from scripted speech and natural from synthetic speech. The researchers employed language isolation probing and centroid proximity analysis to uncover an encoder-dependent trade-off between language discriminability and spontaneity detection, revealing that out-of-domain generalization is more closely linked to the proximity of training systems to unseen TTS embeddings. These findings highlight critical implications for training data selection in developing robust deepfake detection systems in diverse linguistic contexts.
Out-of-domain generalization in speech encoders hinges on their proximity to unseen TTS embeddings rather than their distance from natural speech.
Transformer-based models have shown strong accuracy in distinguishing spontaneous from scripted speech and natural from synthetic speech, but these results are established on a narrow set of well-resourced language benchmarks and have not been extended across Indic languages, nor has embedding geometry been used to explain encoder behaviour or deepfake generalisation failure. We address these gaps by evaluating five frozen transformer encoders, AST, Vaani-FastConformer, Wav2vec2, Whisper and BEATs, across 22 Indic languages, and by conducting a multi-system TTS generalisation experiment across four TTS models. Beyond accuracy, we present language isolation probing and centroid proximity analysis. Probing reveals an encoder-dependent trade-off between language-discriminability and spontaneity detection. Centroid analysis shows that out-of-domain generalisation is predicted by a training system's proximity to unseen TTS embeddings, not its distance from natural speech, a finding with direct implications for training data selection in real-world deepfake detectors.