Search papers, labs, and topics across Lattice.
9
2
10
7
Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model, is presented and the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing is explored.
Achieving a staggering 54.67脳 speedup in text-to-video-audio generation without sacrificing quality could revolutionize real-time multimedia applications.
Robots can now predict not just immediate actions but also the next stages of complex tasks, leading to more efficient manipulation strategies.
Voice-controlled video generation just got a major upgrade with Vidu S1, achieving real-time performance without visual distortion.
Cleaner visual cues can boost multimodal reasoning performance by over 6 points, challenging the notion that simply extending reasoning traces is sufficient.
Streaming video generation can be served 37.5% faster and at 37.2% lower costs with TurboServe's innovative scheduling approach.
Trainable INT8 attention can match full-precision attention during pre-training, but only if you normalize QK and reduce tokens per step.
Achieve an 18.6x speedup in video diffusion models with 97% attention sparsity by learning how to route and combine sparse and linear attention, outperforming heuristic approaches.
SpargeAttention2 achieves 95% attention sparsity in video diffusion models with a 16.2x speedup, proving that trainable sparse attention can significantly outperform training-free methods without sacrificing generation quality.