Search papers, labs, and topics across Lattice.
2
0
4
FlashDiff slashes diffusion model serving latency by up to 97% while boosting throughput by over 2x through intelligent regional execution and scheduling.
LLM serving systems can boost Time-To-First-Token (TTFT) attainment by up to 2.4x simply by prioritizing network flows based on a novel approximation of Least-Laxity-First scheduling.