Search papers, labs, and topics across Lattice.
4
0
4
42
This paper presents Cross-Lingual F5-TTS 2, a simplified framework for transcript-free cross-lingual voice cloning without forced alignment, and makes the syllable-level speaking rate predictor robust to leading and trailing silence through silence-aware augmentation.
An auditable decision record that carries four aligned cues into a late calibration step: a passive detector score, a conditional keyed-probe score on a marked derivative, retrieval support, and a speaker-profile margin, together with explicit disagreement coordinates is answered.
Natural-language instructions can now handle both zero-shot speech synthesis and surgical acoustic editing within a single unified model, operating at 4-step distilled inference speeds without classifier-free guidance.
Real-time speaker-attributed ASR is now feasible with VibeVoice-ASR-Streaming, achieving unmatched accuracy while processing speech on-the-fly.