Efficient Deep Spiking Multi-Layer Perceptrons with Multiplication-Free Inference
arXiv:2306.12465 · doi:10.1109/TNNLS.2024.3394837
Abstract
Advancements in adapting deep convolution architectures for Spiking Neural Networks (SNNs) have significantly enhanced image classification performance and reduced computational burdens. However, the inability of Multiplication-Free Inference (MFI) to align with attention and transformer mechanisms, which are critical to superior performance on high-resolution vision tasks, imposing limitations on these gains. To address this, our research explores a new pathway, drawing inspiration from the progress made in Multi-Layer Perceptrons (MLPs). We propose an innovative spiking MLP architecture that uses batch normalization to retain MFI compatibility and introducing a spiking patch encoding layer to enhance local feature extraction capabilities. As a result, we establish an efficient multi-stage spiking MLP network that blends effectively global receptive fields with local feature extraction for comprehensive spike-based computation. Without relying on pre-training or sophisticated SNN training techniques, our network secures a top-1 accuracy of 66.39% on the ImageNet-1K dataset, surpassing the directly trained spiking ResNet-34 by 2.67%. Furthermore, we curtail computational costs, model parameters, and simulation steps. An expanded version of our network compares with the performance of the spiking VGG-16 network with a 71.64% top-1 accuracy, all while operating with a model capacity 2.1 times smaller. Our findings highlight the potential of our deep SNN architecture in effectively integrating global and local learning abilities. Interestingly, the trained receptive field in our network mirrors the activity patterns of cortical cells. Source codes are publicly accessible at https://github.com/EMI-Group/mixer-snn.
IEEE TNNLS
References in corpus (16)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MLP-Mixer: An all-MLP Architecture for Vision
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- Deep Residual Learning in Spiking Neural Networks
- Spikformer: When Spiking Neural Network Meets Transformer
- Temporal Efficient Training of Spiking Neural Network via Gradient Re-weighting
- Constructing Accurate and Efficient Deep Spiking Neural Networks with Double-threshold and Augmented Schemes
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
- AutoSNN: Towards Energy-Efficient Spiking Neural Networks
- Cortical oscillations implement a backbone for sampling-based computation in spiking neural networks
- DynaMixer: A Vision MLP Architecture with Dynamic Mixing
- Accurate and Efficient Event-based Semantic Segmentation Using Adaptive Spiking Encoder-Decoder Network
- SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation
- SpikingJelly: An open-source machine learning infrastructure platform for spike-based intelligence