Search papers, labs, and topics across Lattice.
6
0
7
9
This paper presents Cross-Lingual F5-TTS 2, a simplified framework for transcript-free cross-lingual voice cloning without forced alignment, and makes the syllable-level speaking rate predictor robust to leading and trailing silence through silence-aware augmentation.
Natural-language instructions can now handle both zero-shot speech synthesis and surgical acoustic editing within a single unified model, operating at 4-step distilled inference speeds without classifier-free guidance.
AgenticASR revolutionizes speech recognition by enabling real-time, intent-preserving transcription that adapts as speech evolves, outperforming traditional methods.
The shift from basic to high-level semantic intelligence in AI mirrors human cognitive development, revealing critical insights for future advancements in machine understanding and generation.
Real-time multilingual speech translation can now maintain speaker identity and quality even in complex, multi-speaker scenarios.
Achieve human-like full-duplex voice interactions with SoulX-Duplug, a plug-and-play module that slashes latency and improves turn management by acting as a semantic VAD.