Search papers, labs, and topics across Lattice.
2
0
5
4
Jet-Long achieves up to 1.39x throughput improvements while maintaining accuracy across long-context tasks, setting a new standard for zero-shot context extension in LLMs.
K-means, previously relegated to offline processing, gets a 17.9x speed boost on modern GPUs thanks to Flash-KMeans' clever IO and contention optimizations.