Search papers, labs, and topics across Lattice.
This paper explores the limitations of single-channel source separation by analyzing the geometric properties of time-frequency masking, revealing that traditional oracle masks do not achieve optimal performance. The authors introduce a structured approach that utilizes a chain of nested classes based on assumptions about the source prior, demonstrating that the phase component significantly influences the estimation accuracy. Their method achieves a mean-square estimate that outperforms existing techniques on the MUSDB18 dataset, achieving a notable 11.44 dB improvement under the per-frame ceiling while attributing a significant portion of the error to predicted variance.
Traditional single-channel separation methods fall short of optimal performance, but this new geometric approach reveals a pathway to significantly reduce estimation error by addressing phase discrepancies.
Most single-channel separators estimate a source by applying a real gain to the mixture in each time-frequency bin. The optimum of that format, which the oracle masks used as bounds do not attain, is the orthogonal projection of the source onto the line spanned by the mixture, its residual set by the angle between them. Locating an estimator reduces to the block structure of a real-linear operator on stacked spectra, giving a chain of four nested classes whose three larger terms match three assumptions on the prior: zero means, circularity and absence of inter-frequency coupling. Held fixed the chain is a cascade of four orthogonal projections; refitted per frame it collapses onto its first term, attributing the whole residual to one missing real parameter per bin, the phase. When the phase posterior is symmetric about the mixture direction, the minimum mean-square estimate falls back onto the line, with gain the posterior mean of the oracle gain and excess error its variance. On MUSDB18 a posterior mean under a non-circular Gaussian-mixture prior leaves the class yet stays 11.44 dB under the per-frame ceiling, which four times as many components and 7.5x the data do not close; a closed-form gate attributes some 70% of it, in decibels, to the predicted variance. The widest fixed class stays 6.70 dB under the same ceiling. Leaving the class and minimising squared error are conflicting requests: the barrier lies in the criterion rather than in the prior.