Search papers, labs, and topics across Lattice.
4
0
5
4
FireRedAudio achieves state-of-the-art performance in audio understanding and generation by leveraging decoupled continuous representations, setting a new standard for unified audio-language models.
Emotion preference models can be dramatically improved by addressing both data sparsity and model bias, leading to more accurate emotional assessments in multimodal contexts.
Semantically enriched speech representations in FireRedTTS3 lead to unprecedented stability and fidelity in voice cloning and editing tasks.
LLMs writing long stories frequently contradict themselves on basic facts and timelines, especially in the middle of the narrative, highlighting a critical weakness in long-form generation.