Search papers, labs, and topics across Lattice.
Vidu S1 is a groundbreaking real-time interactive video generation model that allows users to control video content dynamically through voice commands. Utilizing TurboDiffusion and TurboServe, it generates high-quality 540p videos at up to 42 FPS on standard consumer GPUs, enabling infinite-length video creation without visual artifacts. Experimental results indicate that Vidu S1 outperforms existing models across all evaluated metrics while maintaining real-time performance, making it a significant advancement in interactive media technology.
Voice-controlled video generation just got a major upgrade with Vidu S1, achieving real-time performance without visual distortion.
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions. Vidu S1 supports infinite-length real-time video generation without blurring, drift, or visual distortion. Built with TurboDiffusion and TurboServe, Vidu S1 outputs 540p real-time videos at up to 42 FPS on regular consumer GPUs. Users can upload custom images of real people, anime, and pets, and choose different voice tones for personalized experiences. Experiments show that Vidu S1 achieves the best performance across all test metrics while fully meeting real-time inference requirements. A playable online demo is available at https://vidu.com/vidu-stream.