Search papers, labs, and topics across Lattice.
2
0
3
PLoRA slashes decode latency for multi-LoRA serving by over 6x while using pooled memory and near-data processing, revolutionizing how we deploy specialized AI models.
Forget GPU-centric designs: AMMA slashes attention latency by 15x and energy consumption by 7x with a memory-centric architecture for long-context LLMs.