Search papers, labs, and topics across Lattice.
3
0
6
9
Task-Agnostic Pretraining enables VLA models to achieve expert-level performance with orders of magnitude less labeled data, revolutionizing the scalability of embodied AI.
A purely Transformer-based audio tokenizer, pre-trained on 3M hours of data, leapfrogs existing codecs and even enables a fully autoregressive TTS model to outperform cascaded systems.
Open-source MOVA lets you generate synchronized, high-quality video and audio—including realistic lip sync—without relying on closed-source systems.