Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to memory partitioning in Spiking Neural Networks (SNNs) that optimizes power consumption by leveraging the varying firing rates of neurons. By allocating synaptic weights to multiple non-uniformly sized memory banks based on their access frequency, the proposed architecture significantly reduces power usage while minimizing area overhead. Experimental results demonstrate a power reduction of up to 61% in synaptic weight memory access compared to conventional designs, with a 2.1脳 lower area overhead than traditional uniformly partitioned memory banks.
Allocating synaptic weights based on neuron firing rates can slash power consumption by over 60% in Spiking Neural Networks without significant area costs.
Spiking Neural Networks (SNNs) naturally excel in processing temporally rich and sparse data. However, because of their time-stepped processing, memory access, specifically to synaptic weights stored in SRAM (static random-access memory), tends to dominate total power consumption. To address this issue, without incurring a large area overhead, we propose to leverage the greatly varying average firing rate of neurons in the network to efficiently allocate synaptic weights to an on-chip memory consisting of multiple non-uniformly sized memory banks. By assigning weights of frequently firing neurons to shallow, low-access cost memory and less actively accessed weights to deeper, high-density memories, the average power consumption of the synaptic weight memory is decreased without incurring a large area overhead. To benchmark our proposed architecture and find optimal configurations of memory arrangements, we perform an automatic exploration based on application requirements and hardware constraints. For memory designs synthesized in 28-nm CMOS technology, we show that our architecture can achieve a synaptic weight memory access power reduction of up to 61\% compared to a conventional design, with a 2.1$\times$ lower area overhead, as compared to a traditional uniformly partitioned memory bank that achieves a comparable reduction.