Search papers, labs, and topics across Lattice.
University of Minnesota
2
0
5
2
SparseDitto achieves up to 146.61x speedup for sparse matrix operations by dynamically customizing GPU kernels based on input patterns.
Forget hand-tuning CUDA: a multi-agent system now automates end-to-end GPU program generation with near-perfect success and significant speedups.