Search papers, labs, and topics across Lattice.
2
0
4
FourTune slashes memory overhead by 2.25x while matching the performance of full-precision fine-tuning in diffusion models.
K-means, previously relegated to offline processing, gets a 17.9x speed boost on modern GPUs thanks to Flash-KMeans' clever IO and contention optimizations.