CUHKMar 5, 2026arXiv:2603.05091

Voice Timbre Attribute Detection with Compact and Interpretable Training-Free Acoustic Parameters

Aemon Yat Fei Chiu, Yujia Xiao, Qiuqiang Kong, Tan Lee

AI Summary

This paper explores the use of a compact, training-free acoustic parameter set for voice timbre attribute detection (vTAD), a task that measures the relative intensity of timbre attributes between speech utterances. The proposed parameter set captures crucial acoustic measures and their temporal dynamics without requiring any trainable parameters. Experiments demonstrate that this simple acoustic parameter set outperforms conventional cepstral features and supervised DNN embeddings, approaching the performance of state-of-the-art self-supervised models while offering interpretability and negligible computation.

Key Contribution

Surprisingly, a compact, training-free set of acoustic parameters rivals DNN embeddings and approaches self-supervised models in voice timbre attribute detection, offering interpretability and efficiency.

Abstract

Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently complex component of speech perception. While deep neural network (DNN) embeddings perform well in speaker modelling, they often act as black-box representations with limited physical interpretability and high computational cost. In this work, a compact acoustic parameter set is investigated for vTAD. The set captures important acoustic measures and their temporal dynamics which are found to be crucial in the task. Despite its simplicity, the acoustic parameter set is competitive, outperforming conventional cepstral features and supervised DNN embeddings, and approaching state-of-the-art self-supervised models. Importantly, the studied set require no trainable parameters, incur negligible computation, and offer explicit interpretability for analysing physical traits behind human timbre perception.

Interpretability & Mechanistic Interp Speech & Audio

Citation Metrics

Citations0

Influential citations0

References39

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Voice Timbre Attribute Detection with Compact and Interpretable Training-Free Acoustic Parameters

Related Papers