Search papers, labs, and topics across Lattice.
This paper introduces a training-free dynamic clustering method for aligning cross-segment permutations in long speech separation by leveraging speaker embedding reference pools. By utilizing cosine similarity to predict permutations and updating reference pools with the most representative embeddings, the approach enhances the stitching process of independently processed segments. The method outperforms existing techniques, especially in challenging scenarios with sparse utterances and unknown speaker counts, making it a robust solution for long speech separation tasks.
A training-free dynamic clustering method significantly improves long speech separation performance, especially in sparse scenarios with unknown speaker counts.
Long speech separation typically employs a segment-separation-stitch paradigm where recordings are divided into short segments, processed independently, and stitched together. Its challenge lies in predicting cross-segment permutations. This paper proposes a training-free dynamic clustering approach for cross-segment permutation alignment using speaker embedding reference pools. The method predicts the permutation using the cosine similarity between current segment embeddings and the reference pools. The approach updates reference pools by retaining the most representative speaker embeddings based on their overall cosine similarity with existing references. As a plug-and-play post-processing module compatible with existing separation models, the proposed method demonstrates superior performance compared to existing methods on dense and sparse long speech scenarios, particularly in challenging sparse scenarios with extended utterance gaps, and further shows robustness to speaker count estimation errors in unknown speaker count scenarios.