Search papers, labs, and topics across Lattice.
University of Minnesota
4
0
7
5
SparseDitto achieves up to 146.61x speedup for sparse matrix operations by dynamically customizing GPU kernels based on input patterns.
LACE transforms the RISC-V instruction extension process, achieving a 72.8% accuracy in generating ISAX implementations from natural language intents, a leap from near-zero accuracy.
Zeroth-order fine-tuning can be sped up by over 8x by reframing it as an inference workload and executing it within a serving runtime.
Forget hand-tuning CUDA: a multi-agent system now automates end-to-end GPU program generation with near-perfect success and significant speedups.