Search papers, labs, and topics across Lattice.
This paper explores a novel approach to neuromorphic speech recognition by developing a high-level, programmable encoder that converts audio data into spikes for processing with Spiking Neural Networks (SNNs). By optimizing both the encoder and classifier simultaneously, the authors achieve a remarkable classification accuracy of 99.77% on the Heidelberg Digits dataset while significantly reducing energy costs during learning and inference. This work not only establishes a new benchmark in neuromorphic audio processing but also addresses the challenges posed by the scarcity of native neuromorphic datasets.
Achieving 99.77% accuracy in neuromorphic speech recognition while slashing energy costs could redefine efficiency standards in AI processing.
Obtaining data from neuromorphic sensors and processing it with Spiking Neural Networks is a promising solution to lower the energy cost of artificial intelligence. The current rarity of natively neuromorphic datasets promotes the development of software tools to translate input sensory data into spikes. However, highly bio-mimetic simulators can be challenging to implement on digital hardware. In this work, we evaluate the neuromorphic encoding and subsequent classification of audio into spikes using a non-learnable, high-level, programmable encoder targeting hardware implementation on FPGA. We quantify the pipeline's efficiency with hardware-agnostic metrics based on the quantitative spiking activity. Our study focuses on the simultaneous optimisation of encoder and classifier: the first provides efficient and informative data so that the latter achieves a better performance with an overall lower energy cost at learning and inference. This work introduces the first end-to-end neuromorphic spike-encoding and evaluation of the TIMIT dataset. Our simple feedforward network reaches a classification accuracy of 99.77% on a spike-encoded Heidelberg Digits, overcoming the neuromorphic state of the art on this benchmark dataset.