Search papers, labs, and topics across Lattice.
4
0
6
4
Voice-controlled video generation just got a major upgrade with Vidu S1, achieving real-time performance without visual distortion.
Only a subset of design interactions in heterogeneous LLM inference are binding constraints, revealing critical insights for optimizing deployment strategies.
Streaming video generation can be served 37.5% faster and at 37.2% lower costs with TurboServe's innovative scheduling approach.
LLM serving can be sped up by 50% on average by dynamically adapting model deployments to match the changing mix of request types.