Search papers, labs, and topics across Lattice.
8
2
10
11
Voice-controlled video generation just got a major upgrade with Vidu S1, achieving real-time performance without visual distortion.
Teacher-forcing consistency models can accelerate autoregressive video generation by ten times, revolutionizing the training landscape for streaming applications.
Streaming video generation can be served 37.5% faster and at 37.2% lower costs with TurboServe's innovative scheduling approach.
LLMs can generate GPU kernels, but they're surprisingly bad at it: 72% of fusion tasks fail across all methods, and nearly half of the "correct" kernels are actually slower than PyTorch.
Video diffusion models can be aggressively quantized down to 6-bit precision with minimal quality loss by dynamically adapting the bit-width of each layer based on its temporal stability.
Ditch the stochasticity: Deterministic pruning slashes LLM size with minimal performance loss, outperforming stochastic methods and accelerating inference.
Achieve an 18.6x speedup in video diffusion models with 97% attention sparsity by learning how to route and combine sparse and linear attention, outperforming heuristic approaches.
SpargeAttention2 achieves 95% attention sparsity in video diffusion models with a 16.2x speedup, proving that trainable sparse attention can significantly outperform training-free methods without sacrificing generation quality.