Search papers, labs, and topics across Lattice.
Yale University
1
0
2
Celty achieves up to 5.3x speedup in LLM inference by harnessing dual-sparsity, revolutionizing how we approach GPU kernel design for sparse workloads.