Search papers, labs, and topics across Lattice.
The Qwen-Audio-3.0-Gen-Preview introduces a unified non-autoregressive framework that leverages a Diffusion Transformer and a shared variational autoencoder to generate complete mixed waveforms from heterogeneous audio components. This approach enhances prompt conversion into structured temporal records, allowing for improved performance in multi-speaker and rich-timeline benchmarks, particularly in speaker similarity and temporal localization. The results indicate that this model can effectively generate temporally structured audio without relying on task-specific branches, showcasing its versatility across various audio tasks.
Unified generation of temporally structured audio achieves unprecedented speaker similarity and cross-turn consistency without task-specific branches.
Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion Transformer (DiT) and a shared variational autoencoder (VAE) to generate the complete mixed waveform. Prompt enhancement converts free-form requests into structured temporal records that are rendered as textual conditions, while a two-stage data curriculum and semantic conditional views train the proposed model to use these conditions across standalone and mixed-scene audio. A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences and incorporates semantic supervision, providing one representation for heterogeneous audio. On the public reference-conditioned benchmark, speaker similarity is the proposed model's clearest strength across all three subsets. Across the multi-speaker and rich-timeline benchmarks, its clearest comparative strengths are cross-turn consistency in both languages and temporal localization, respectively. On AudioCaps, its advantages are concentrated in evaluations using large audio-language models and AudioBox. These results demonstrate the potential of unified generation for temporally structured audio without task-specific branches.