Search papers, labs, and topics across Lattice.
Qwen-Audio-3.0-TTS is a cutting-edge speech synthesis system that integrates a low-frame-rate speech tokenizer with a five-stage progressive training paradigm to enhance content consistency, speaker similarity, and audio quality across multiple languages. The model allows for versatile control through natural language instructions and inline tags, while also demonstrating robustness against noisy and reverberant inputs. Achieving state-of-the-art results in various evaluations, including long-form synthesis and acoustic robustness, it establishes itself as a leading solution in production-level speech synthesis.
Achieving state-of-the-art performance in speech synthesis, Qwen-Audio-3.0-TTS excels in multilingual support and robustness against challenging audio conditions.
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated language model (LM) and flow-matching model (FM) optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.