Search papers, labs, and topics across Lattice.
This paper introduces Chiplet-Contiguous Layout, a novel global memory layout that enhances locality-aware data placement for general matrix multiplication (GEMM) on multi-chiplet GPUs. By storing chiplet-local data contiguously, the method achieves compatibility with page-granularity data interleaving, addressing the challenge of varying optimal granularity across workloads. The results show significant reductions in remote high-bandwidth memory (HBM) traffic, achieving up to 24.7x lower traffic for Qwen 3 30B and 19.2x for Llama 3.1 70B compared to traditional interleaving methods.
Achieving up to 24.7x reduction in remote HBM traffic for LLM GEMM tasks could redefine efficiency standards in multi-chiplet GPU architectures.
Multi-chiplet GPUs scale compute throughput and high-bandwidth memory (HBM) capacity, but their non-uniform memory system makes locality between chiplets and their data critical to the GPU's performance and energy efficiency. Locality-aware scheduling and data placement identify which data should reside near each chiplet. However, in general matrix multiplication (GEMM), locality-aware data placement often becomes incompatible with a fixed page-granularity data interleaving, since the optimal granularity for mapping data across chiplets varies widely across workloads. We propose Chiplet-Contiguous Layout, a global memory layout that stores chiplet-local data contiguously. Chiplet-Contiguous Layout enables locality-aware placement compatible with page-granularity placement across diverse large language model (LLM) GEMM shapes, without changes to the operating system or hardware. On representative LLM inference and training GEMMs from Qwen 3 30B and Llama 3.1 70B, Chiplet-Contiguous Layout on average reduces remote HBM traffic by 24.7x on Qwen and 19.2x on Llama over 4KB interleaving, and by 4.1x and 2.1x over coarse locality-aware placement.