VMamba: Visual State Space Model
arXiv:2401.10166
Abstract
Designing computationally efficient network architectures remains an ongoing necessity in computer vision. In this paper, we adapt Mamba, a state-space language model, into VMamba, a vision backbone with linear time complexity. At the core of VMamba is a stack of Visual State-Space (VSS) blocks with the 2D Selective Scan (SS2D) module. By traversing along four scanning routes, SS2D bridges the gap between the ordered nature of 1D selective scan and the non-sequential structure of 2D vision data, which facilitates the collection of contextual information from various sources and perspectives. Based on the VSS blocks, we develop a family of VMamba architectures and accelerate them through a succession of architectural and implementation enhancements. Extensive experiments demonstrate VMamba's promising performance across diverse visual perception tasks, highlighting its superior input scaling efficiency compared to existing benchmark models. Source code is available at https://github.com/MzeroMiko/VMamba.
33 pages, 14 figures, 15 tables. NeurIPS 2024 spotlight
Cited by in corpus (8)
- H-vmunet: High-order Vision Mamba UNet for Medical Image Segmentation
- UltraLight VM-UNet: Parallel Vision Mamba Significantly Reduces Parameters for Skin Lesion Segmentation
- Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
- Medical Image Classification with KAN-Integrated Transformers and Dilated Neighborhood Attention
- EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention
- Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
- MedVKAN: Efficient Feature Extraction with Mamba and KAN for Medical Image Segmentation
- Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation