Search papers, labs, and topics across Lattice.
This paper introduces Lonic, a novel algorithm-hardware co-design aimed at enhancing the energy efficiency of fully local online training for spiking neural networks (SNNs) using INT4 precision. By integrating a low-precision training algorithm with innovative hardware components, including reconfigurable multiplier-free integer processing elements and a dual-optimization zero-gating strategy, Lonic achieves significant improvements in both energy and area efficiency. The results demonstrate that Lonic outperforms existing hardware solutions, achieving up to 66.28x energy efficiency gains compared to Nvidia V100 GPUs while maintaining competitive training speeds.
Lonic achieves up to 66.28x energy efficiency improvements over leading GPUs, revolutionizing the training landscape for spiking neural networks.
Spiking neural networks (SNNs) have recently attracted increasing attention as an energy-efficient learning paradigm. Existing works also propose temporally and fully local online SNN training algorithms to address memory and computation overhead. However, they do not consider whether the algorithmic advantages can be effectively translated into real-device efficiency. To address this challenge, we present Lonic, an algorithm-hardware co-design for energy-efficient and scalable fully local online supervised SNN learning. On the algorithm side, we implement an INT4 low-precision training algorithm for fully local online SNN learning while maintaining accuracy. On the hardware side, to leverage the benefits of the proposed algorithm, we introduce reconfigurable multiplier-free integer PE arrays, dual-optimization zero-gating strategy, temporal prefix-accelerated local learning dataflow, and low-precision weight movement to significantly improve training efficiency. Compared to Apple M4 and Nvidia V100 GPUs, Lonic achieves average energy efficiency improvements of 17.44x and 66.28x, respectively, along with speedups of 3.25x and 1.02x, respectively. Moreover, Lonic achieves 15.95x (14.64x) and 1.52x (7.28x) energy efficiency (area efficiency) over ASIC TPU-like and H2Learn accelerators, respectively. The code for Lonic is available at https://github.com/peilin-chen/Lonic.