Search papers, labs, and topics across Lattice.
7
0
9
3
GROW achieves a 22.7% reduction in word error rate while accelerating training by 2.9x, redefining efficiency in TTS reinforcement learning.
Runtime load balancing in FVAttn slashes attention processing time by over four times, transforming video generation efficiency.
OmniAgent not only outperforms larger models but also scales performance with reasoning turns, revolutionizing how we approach video understanding.
Continuous-target modeling reveals a shared semantic mapping for ASR and S2TT, challenging conventional views on their error sources.
Current audio editing models are failing spectacularly, with an Exact Match Rate below 5% in complex tasks, exposing a critical need for improvement.
Real-time audio interaction is now possible with a unified model that not only performs traditional tasks but also proactively responds to audio stimuli.
Online reinforcement learning with large audio language model rewards catapults text-to-audio generation to a new state-of-the-art, even with a relatively small 470M parameter model.