Search papers, labs, and topics across Lattice.
3
0
4
5
FireRedAudio achieves leading performance in audio understanding and speech generation by decoupling input representations, marking a significant advancement in unified audio-language modeling.
Semantically enriched speech representations in FireRedTTS3 lead to unprecedented stability and fidelity in voice cloning and editing tasks.
A fully open-source speech understanding model, OSUM-Pangu, proves that competitive performance is achievable on non-CUDA hardware, challenging the dominance of GPU-centric ecosystems.