Search papers, labs, and topics across Lattice.
2
0
4
2
LLMs struggle to match expert-level performance in GPU communication tasks, with top models achieving only 30.7% success in generating efficient code.
Diffusion language models can achieve faster convergence and improved accuracy simply by swapping token-choice routing for expert-choice routing, and further benefit from allocating more compute to early denoising steps.