Search papers, labs, and topics across Lattice.
4
17
5
34
InvFlowFD achieves reference-free music quality assessment, aligning closely with human perception while eliminating the need for background data.
Non-verbal vocalizations can be effectively modeled in ASR, improving recognition of rare events without sacrificing lexical accuracy.
Ditch the pre-trained models: PAST directly learns speech tokens from phonetic data, outperforming existing methods in representation and reconstruction.
Edit the bassline, drums, or other instruments of any song with this new open-source multi-stem music generation model.