Abstract
we compared three computationally simple methods for encoding short speech recordings as spike sequences for real-time classification with a multilayer spiking neural network. The methods included integral coding of spectral features, stream spectral encoding, and direct encoding of local signal dynamics. All three methods used the same network of Leaky Integrate-and-Fire neurons trained by surrogate-gradient backpropagation. Experiments on the Google Speech Commands v0.01 (GSC) and Heidelberg Digits (HD) datasets showed that integral rate coding achieved the highest accuracy on GSC (83.1%), whereas all three methods performed comparably on HD. Stream spectral encoding required only four simulation time steps per spectrogram frame, substantially reducing response latency. Speech2Spikes achieved 80.1% accuracy on GSC, only 3% below integral rate coding, while using a more compact input representation. The results provide a basis for selecting spike encoding methods under different computational and latency constraints.

This work is licensed under a Creative Commons Attribution 4.0 International License.
