Search papers, labs, and topics across Lattice.
KAIST
2
0
4
5
NELSSA achieves a staggering 5.5x increase in decode throughput for mixed-length LLM workloads by intelligently integrating GPUs with PNM accelerators.
Stop guessing how your disaggregated LLM serving setup will perform: this simulator nails runtime interactions between hardware and software with near-perfect accuracy.