Efficiently Modeling Long Sequences with Structured State Spaces
arXiv:2111.00396
Abstract
A central goal of sequence modeling is designing a single principled model that can address sequence data across a range of modalities and tasks, particularly on long-range dependencies. Although conventional models including RNNs, CNNs, and Transformers have specialized variants for capturing long dependencies, they still struggle to scale to very long sequences of or more steps. A promising recent approach proposed modeling sequences by simulating the fundamental state space model (SSM) \( x'(t) = Ax(t) + Bu(t), y(t) = Cx(t) + Du(t) \), and showed that for appropriate choices of the state matrix \( A \), this system could handle long-range dependencies mathematically and empirically. However, this method has prohibitive computation and memory requirements, rendering it infeasible as a general sequence modeling solution. We propose the Structured State Space sequence model (S4) based on a new parameterization for the SSM, and show that it can be computed much more efficiently than prior approaches while preserving their theoretical strengths. Our technique involves conditioning \( A \) with a low-rank correction, allowing it to be diagonalized stably and reducing the SSM to the well-studied computation of a Cauchy kernel. S4 achieves strong empirical results across a diverse range of established benchmarks, including (i) 91\% accuracy on sequential CIFAR-10 with no data augmentation or auxiliary losses, on par with a larger 2-D ResNet, (ii) substantially closing the gap to Transformers on image and language modeling tasks, while performing generation faster (iii) SoTA on every task from the Long Range Arena benchmark, including solving the challenging Path-X task of length 16k that all prior work fails on, while being as efficient as all competitors.
ICLR 2022 (Outstanding Paper HM)
References in corpus (14)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- On the difficulty of training Recurrent Neural Networks
- MLP-Mixer: An all-MLP Architecture for Vision
- Generating Long Sequences with Sparse Transformers
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- GRU-ODE-Bayes: Continuous modeling of sporadically-observed time series
- Trellis Networks for Sequence Modeling
- Adversarial Audio Synthesis
- Cheap Orthogonal Constraints in Neural Networks: A Simple Parametrization of the Orthogonal and Unitary Group
- Lipschitz Recurrent Neural Networks
- CKConv: Continuous Kernel Convolution For Sequential Data
- Parallelizing Legendre Memory Unit Training
Cited by in corpus (16)
- TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting
- H-vmunet: High-order Vision Mamba UNet for Medical Image Segmentation
- Transformers in Healthcare: A Survey
- Generative AI for Synthetic Data Across Multiple Medical Modalities: A Systematic Review of Recent Developments and Challenges
- Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
- SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
- GenHPF: General Healthcare Predictive Framework with Multi-task Multi-source Learning
- CT-Mamba: A Hybrid Convolutional State Space Model for Low-Dose CT Denoising
- MambaMIM: Pre-training Mamba with State Space Token Interpolation and its Application to Medical Image Segmentation
- A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint Models
- Enhancing clinical decision support with physiological waveforms -- a multimodal benchmark in emergency care
- Why does my medical AI look at pictures of birds? Exploring the efficacy of transfer learning across domain boundaries
- BarcodeBERT: Transformers for Biodiversity Analysis
- Predictive Modeling in the Reservoir Kernel Motif Space
- Infrastructure-based End-to-End Learning and Prevention of Driver Failure
- State-Space Model Inspired Multiple-Input Multiple-Output Spiking Neurons