MambaMIM: Pre-training Mamba with State Space Token Interpolation and its Application to Medical Image Segmentation
arXiv:2408.08070 · doi:10.1016/j.media.2025.103606
Abstract
Recently, the state space model Mamba has demonstrated efficient long-sequence modeling capabilities, particularly for addressing long-sequence visual tasks in 3D medical imaging. However, existing generative self-supervised learning methods have not yet fully unleashed Mamba's potential for handling long-range dependencies because they overlook the inherent causal properties of state space sequences in masked modeling. To address this challenge, we propose a general-purpose pre-training framework called MambaMIM, a masked image modeling method based on a novel TOKen-Interpolation strategy (TOKI) for the selective structure state space sequence, which learns causal relationships of state space within the masked sequence. Further, MambaMIM introduces a bottom-up 3D hybrid masking strategy to maintain a masking consistency across different architectures and can be used on any single or hybrid Mamba architecture to enhance its multi-scale and long-range representation capability. We pre-train MambaMIM on a large-scale dataset of 6.8K CT scans and evaluate its performance across eight public medical segmentation benchmarks. Extensive downstream experiments reveal the feasibility and advancement of using Mamba for medical image pre-training. In particular, when we apply the MambaMIM to a customized architecture that hybridizes MedNeXt and Vision Mamba, we consistently obtain the state-of-the-art segmentation performance. The code is available at: https://github.com/FengheTan9/MambaMIM.
Accepted by Medical Image Analysis. Code: https://github.com/FengheTan9/MambaMIM
References in corpus (31)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Bootstrap your own latent: A new approach to self-supervised Learning
- TotalSegmentator: robust segmentation of 104 anatomical structures in CT images
- The Medical Segmentation Decathlon
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- BEiT: BERT Pre-Training of Image Transformers
- Efficiently Modeling Long Sequences with Structured State Spaces
- Submanifold Sparse Convolutional Networks
- U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation
- iBOT: Image BERT Pre-Training with Online Tokenizer
- AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation
- WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image
- Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
- LightM-UNet: Mamba Assists in Lightweight UNet for Medical Image Segmentation
- Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling
- The KiTS21 Challenge: Automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase CT
- Vision Mamba: A Comprehensive Survey and Taxonomy
- LocalMamba: Visual State Space Model with Windowed Selective Scan
- PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
- How Mask Matters: Towards Theoretical Understandings of Masked Autoencoders
- Unleashing the Strengths of Unlabeled Data in Pan-cancer Abdominal Organ Quantification: the FLARE22 Challenge
- MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection
- The Hidden Uniform Cluster Prior in Self-Supervised Learning
- MambaMorph: a Mamba-based Framework for Medical MR-CT Deformable Registration
- MambaMIR: An Arbitrary-Masked Mamba for Joint Medical Image Reconstruction and Uncertainty Estimation
- A Comprehensive Survey of Mamba Architectures for Medical Image Analysis: Classification, Segmentation, Restoration and Beyond
- Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
- Scalable Visual State Space Model with Fractal Scanning
- GroupMamba: Efficient Group-Based Visual State Space Model
- MedSegMamba: 3D CNN-Mamba Hybrid Architecture for Brain Segmentation
- Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion