Search papers, labs, and topics across Lattice.
This report details the winning approach to the AES AIMLA 2025 Challenge, focusing on querying sound effects through vocal imitation. The authors explore two fine-tuning strategies: one utilizing contrastive learning with a frozen CED encoder, and the other employing joint contrastive-triplet learning with semi-hard negatives via a MobileNetV3 encoder. The findings highlight the effectiveness of these strategies in improving sound effect retrieval accuracy, which is crucial for applications in audio processing and machine learning.
Fine-tuning with semi-hard negatives can significantly enhance the accuracy of sound effect retrieval through vocal imitation.
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.