Search papers, labs, and topics across Lattice.
8
0
8
12
MOSS-VL achieves a staggering 66.0 average score in proactive alerting, far surpassing the best baseline by nearly 30 points.
Fine-grained cross-modal alignment in audio-video generation can dramatically enhance synchronization and quality, as shown by OmniVAE's innovative training approach.
OmniAct achieves unprecedented levels of physical autonomy, outperforming existing systems by seamlessly integrating multimodal planning and adaptive memory management.
MOSS-Audio achieves state-of-the-art performance in audio understanding tasks by effectively integrating temporal cues and deep acoustic features, setting a new benchmark for audio-language models.
Cinematic speech data unlocks more realistic and controllable voice generation from natural language descriptions.
Achieve controllable and scalable speech generation with MOSS-TTS, enabling zero-shot voice cloning and long-form synthesis.
A purely Transformer-based audio tokenizer, pre-trained on 3M hours of data, leapfrogs existing codecs and even enables a fully autoregressive TTS model to outperform cascaded systems.
Open-source MOVA lets you generate synchronized, high-quality video and audio—including realistic lip sync—without relying on closed-source systems.