Mamba: Linear-Time Sequence Modeling with Selective State Spaces
arXiv:2312.00752
Abstract
Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers' computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5 higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pretraining and downstream evaluation.
Cited by in corpus (35)
- MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image Classification
- H-vmunet: High-order Vision Mamba UNet for Medical Image Segmentation
- UltraLight VM-UNet: Parallel Vision Mamba Significantly Reduces Parameters for Skin Lesion Segmentation
- Materials science in the era of large language models: a perspective
- Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
- Spatial and Spatial-Spectral Morphological Mamba for Hyperspectral Image Classification
- WaveMamba: Spatial-Spectral Wavelet Mamba for Hyperspectral Image Classification
- FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
- Deep learning based infrared small object segmentation: Challenges and future directions
- SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
- CT-Mamba: A Hybrid Convolutional State Space Model for Low-Dose CT Denoising
- Reasoning Beyond Limits: Advances and Open Problems for LLMs
- MambaMIM: Pre-training Mamba with State Space Token Interpolation and its Application to Medical Image Segmentation
- MUCM-Net: A Mamba Powered UCM-Net for Skin Lesion Segmentation
- Mapping the Landscape of Generative AI in Network Monitoring and Management
- Towards a Systems Theory of Algorithms
- SparX: A Sparse Cross-Layer Connection Mechanism for Hierarchical Vision Mamba and Transformer Networks
- The Impact of LoRA Adapters on LLMs for Clinical Text Classification Under Computational and Data Constraints
- SSD-TS: Exploring the Potential of Linear State Space Models for Diffusion Models in Time Series Imputation
- Language Modeling on a SpiNNaker 2 Neuromorphic Chip
- Deep Learning for Human Locomotion Analysis in Lower-Limb Exoskeletons: A Comparative Study
- Towards Context-aware Convolutional Network for Image Restoration
- TiM4Rec: An Efficient Sequential Recommendation Model Based on Time-Aware Structured State Space Duality Model
- OSDMamba: Enhancing Oil Spill Detection from Remote Sensing Images Using Selective State Space Model
- Knowledge-data fusion dominated vehicle platoon dynamics modeling and analysis: A physics-encoded deep learning approach
- A Survey of Spatio-Temporal EEG data Analysis: from Models to Applications
- Emergence of Human-Like Attention in Self-Supervised Vision Transformers: an eye-tracking study
- Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
- NeuralPDR: Neural Differential Equations as surrogate models for Photodissociation Regions
- Joint multi-dimensional dynamic attention and transformer for general image restoration
- MambaEviScrib: Mamba and Evidence-Guided Consistency Enhance CNN Robustness for Scribble-Based Weakly Supervised Ultrasound Image Segmentation
- Representation Learning of Point Cloud Upsampling in Global and Local Inputs
- Leveraging Sound Source Trajectories for Universal Sound Separation
- EMK-KEN: A High-Performance Approach for Assessing Knowledge Value in Citation Network
- Learning Feedback Mechanisms for Measurement-Based Variational Quantum State Preparation