Search papers, labs, and topics across Lattice.
2
0
4
8
Achieving up to 47.26x speedup in long-context LLM serving could redefine efficiency benchmarks in AI inference.
Forget slow attention: FlashPrefill achieves a staggering 27x speedup in long-context prefilling by instantly discovering and thresholding sparse attention patterns.