Search papers, labs, and topics across Lattice.
MAIS&NLPR, CASIA
1
0
3
9
Achieving up to 47.26x speedup in long-context LLM serving could redefine efficiency benchmarks in AI inference.