Search papers, labs, and topics across Lattice.
This paper introduces ABI, a novel near-memory GPU architecture designed for improved performance and energy efficiency across diverse workloads including CNNs, GCNs, LLMs, linear programming, and Ising models. ABI integrates a sparsity-aware near-memory circuit and a lightweight softmax circuit, achieving significant energy savings. Experimental results demonstrate 6-16x speedup and 6-13x energy savings compared to the MIAOW GPU, with ABI-enabled MI300 and Blackwell systems showing a 4.5x speedup over baseline systems.
A novel GPU architecture slashes energy consumption by up to 13x and boosts speed by up to 16x across diverse workloads, including LLMs and Ising models.
We present a tightly integrated and unified near-memory GPU architecture that delivers 6 to 16 times speedup and 6 to 13 times energy savings across Convolutional Neural Networks, Graph Convolutional Networks, Linear Programming, Large Language Models, and Ising workloads compared to MIAOW GPU. The design includes a custom sparsity-aware near-memory circuit providing about 1.5 times energy savings, and a lightweight softmax circuit providing about 1.6 times energy savings. The architecture supports reconfigurable compute up to INT16 with dynamic resolution updates and scales efficiently across problem sizes. ABI-enabled MI300 and Blackwell systems achieve about 4.5 times speedup over baseline MI300 and Blackwell.