Search papers, labs, and topics across Lattice.
This paper introduces ExaGEMM, a framework designed for optimizing CPU-driven low-bit General Matrix Multiplication (GEMM) through associative in-register computing. By co-exploring parameterized kernels and lightweight SIMD ISA support, ExaGEMM achieves a remarkable 13.29x latency improvement over traditional software-only approaches while effectively pruning the candidate space by 99.2%. The framework's ability to adapt to mixed-precision workloads highlights its significance in enhancing the efficiency of ML inference on conventional CPUs.
ExaGEMM achieves a staggering 13.29x latency reduction for low-bit GEMM on CPUs, revolutionizing how we approach efficient ML inference.
Low-bit GEMM is increasingly central to efficient ML inference, yet very-low-bit execution remains a poor fit for conventional CPUs. Practical deployment spans fragmented regimes-from 1/2/4-bit weights to varying activation precision-whose feasibility, reuse opportunity, and support cost differ under fixed SIMD and register-file budgets, making lightweight CPU support selection a first-class design problem. We present ExaGEMM, a workload-aware codesign and exploration framework for CPU-native low-bit GEMM via register-resident LUT execution. The key insight is that existing SIMD datapaths already cover table generation and accumulation; the only new hardware is an in-register select/feed mechanism with explicitly modeled cost. ExaGEMM co-explores parameterized kernels and lightweight SIMD ISA support using analytical models of register feasibility, compute cost, memory traffic, and hardware overhead, pruning the candidate space by 99.2% before simulation. It then identifies non-dominated support points and generates ISA specs, gem5 patches, and GEMM kernels for validation. Across representative ML models and CPU targets, ExaGEMM improves latency by 13.29x over software-only baselines, while showing that workload-aware frontier selection is especially important for mixed-precision LLM workloads.