Search papers, labs, and topics across Lattice.
UT Austin
4
0
7
ACID redefines the caching paradigm in video generation, achieving unprecedented speedups while maintaining visual fidelity by dynamically adjusting thresholds based on drift signals.
Sangam slashes latency for diffusion language models by intelligently managing prefill and decode processes, revealing a new paradigm for efficient LLM serving.
RL-ACRGNet achieves significant performance boosts in radiology report generation, setting new benchmarks in both accuracy and clinical relevance.
Stop hand-writing CUDA kernels: CUCo's agent-driven approach co-optimizes computation and communication, slashing LLM training/inference latency by up to 1.57x.