Search papers, labs, and topics across Lattice.
7
0
9
4
Fine-grained cross-modal alignment in audio-video generation can dramatically enhance synchronization and quality, as shown by OmniVAE's innovative training approach.
EAPO revolutionizes LLM reasoning by dynamically integrating prior experiences, leading to consistent performance gains over traditional RLVR methods.
MOSS-Audio achieves state-of-the-art performance in audio understanding tasks by effectively integrating temporal cues and deep acoustic features, setting a new benchmark for audio-language models.
Rigid reward clipping throws away valuable information just beyond the boundary, but a simple stochastic rescue of these signals can substantially boost RLVR performance.
LLM benchmarks are riddled with hidden flaws that even human experts miss, but can be caught with an automated LLM auditor for under $15 per benchmark.
Achieve controllable and scalable speech generation with MOSS-TTS, enabling zero-shot voice cloning and long-form synthesis.
A purely Transformer-based audio tokenizer, pre-trained on 3M hours of data, leapfrogs existing codecs and even enables a fully autoregressive TTS model to outperform cascaded systems.