Search papers, labs, and topics across Lattice.
3
0
6
6
LLMs struggle to match expert-level performance in GPU communication tasks, with top models achieving only 30.7% success in generating efficient code.
K-means, previously relegated to offline processing, gets a 17.9x speed boost on modern GPUs thanks to Flash-KMeans' clever IO and contention optimizations.
Get 2x faster video generation from diffusion transformers without sacrificing quality, thanks to a clever parameter-free error compensation technique.