Search papers, labs, and topics across Lattice.
6
0
7
2
Improved spatial audio generation in real-world speech scenes achieves unprecedented fidelity and accuracy using a visually guided framework.
Pruning can slash the computational cost of text-to-audio models by over 80% without sacrificing quality, but it poses risks to generating critical sound events.
Room embeddings can now be reliably estimated from reverberant speech with a calibrated uncertainty score, enabling selective prediction from just one utterance.
Attention maps from speaker recognition models reveal that GradCAM and LayerCAM excel under different conditions, challenging the one-size-fits-all approach in XAI.
Achieving efficient and precise audio editing with a compact model, this hybrid diffusion transformer outperforms existing methods on complex tasks involving overlapping audio events.
CLAP models suffer from a modality gap that's more than just a "cone effect" – it's a small set of interpretable concept axes dominating similarity, and truncating the rest unlocks near-supervised zero-shot performance.