Search papers, labs, and topics across Lattice.
This paper introduces a hardware-software co-design framework for LLM inference that addresses the challenges of memory bandwidth and computational density through a scalable architecture called LEAP. By integrating specialized processing elements for static weights, dynamic data, and partial result reduction, the framework enhances both throughput and energy efficiency. The proposed solution achieves improvements of at least 1.52 times in throughput and 24.91 times in energy efficiency compared to existing GPU platforms, making it a significant advancement in LLM serving technology.
Achieving over 24 times better energy efficiency in LLM inference could redefine the scalability of AI applications.
LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data generated during run-time. Furthermore, the massive number of parameters in LLM necessitates scale-up architectures where on-chip data movement is often the primary performance bottleneck. This article presents a hardware-software co-design framework that unifies distributed compute, memory, and communication into a seamless processing-communication fabric. On the hardware side, we propose a scalable architecture, named LEAP, that integrates IMC PE, NMC PE, and INC. This allows each hardware layer to execute specialized tasks: IMC for static weights, NMC for dynamic data, and INC for partial result reduction. On the software side, we introduce a partitioning, mapping, and scheduling framework optimized for key metrics in LLM serving, including throughput and latency. To address the distinct computational intensities of the prefill and decode phases, we present a prefill-decode disaggregation approach that dynamically reconfigures PE organizations to maximize resource utilization. Compared to commercial GPU platforms, the proposed architecture provides a throughput and an energy efficiency improvement of $\geq{}1.52\times$ and $24.91\times$, respectively.