Search papers, labs, and topics across Lattice.
3
2
4
4
Despite executing target instructions, LLMs often fail to deliver competitive performance in GPU kernel optimization, especially on complex tasks.
Unlock 2x faster LLM serving and slash warmup times by fusing kernels that gracefully handle dynamic shapes and data dependencies.
LLMs can now autonomously generate and deploy GPU kernels into production LLM engines, thanks to a new standardized framework for benchmarking and integrating these AI-generated kernels.