Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong
2
0
4
LUMEN slashes recovery times in distributed LLM serving by intelligently coordinating load-aware recovery strategies during worker failures.
Get up to 4x faster video generation from diffusion transformers without sacrificing quality, thanks to a new clustering method that slashes attention costs.