Search papers, labs, and topics across Lattice.
University of Surrey
5
0
6
By integrating visual cues with audio processing, this framework achieves unprecedented spatial audio fidelity in complex speech environments.
Pruning can slash the computational cost of text-to-audio models by over 80% without sacrificing quality, but it poses risks to generating critical sound events.
Room embeddings can now be reliably estimated from reverberant speech with a calibrated uncertainty score, enabling selective prediction from just one utterance.
Attention maps from speaker recognition models reveal that GradCAM and LayerCAM excel under different conditions, challenging the one-size-fits-all approach in XAI.
Achieving efficient and precise audio editing with a compact model, this hybrid diffusion transformer outperforms existing methods on complex tasks involving overlapping audio events.