Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
arXiv:2406.14086 · doi:10.1145/3725949.3725967
Abstract
Recent advancements in autoregressive networks with linear complexity have driven significant research progress, demonstrating exceptional performance in large language models. A representative model is the Extended Long Short-Term Memory (xLSTM), which incorporates gating mechanisms and memory structures, performing comparably to Transformer architectures in long-sequence language tasks. Autoregressive networks such as xLSTM can utilize image serialization to extend their application to visual tasks such as classification and segmentation. Although existing studies have demonstrated Vision-LSTM's impressive results in image classification, its performance in image semantic segmentation remains unverified. Our study represents the first attempt to evaluate the effectiveness of Vision-LSTM in the semantic segmentation of remotely sensed images. This evaluation is based on a specifically designed encoder-decoder architecture named Seg-LSTM, and comparisons with state-of-the-art segmentation networks. Our study found that Vision-LSTM's performance in semantic segmentation was limited and generally inferior to Vision-Transformers-based and Vision-Mamba-based models in most comparative tests. Future research directions for enhancing Vision-LSTM are recommended. The source code is available from https://github.com/zhuqinfeng1999/Seg-LSTM.
References in corpus (14)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
- VMamba: Visual State Space Model
- VM-UNet: Vision Mamba UNet for Medical Image Segmentation
- LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation
- Advancements in Point Cloud Data Augmentation for Deep Learning: A Survey
- xLSTM: Extended Long Short-Term Memory
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
- Visual Mamba: A Survey and New Outlooks
- MambaOut: Do We Really Need Mamba for Vision?
- SBSS: Stacking-Based Semantic Segmentation Framework for Very High Resolution Remote Sensing Image
- Vision-LSTM: xLSTM as Generic Vision Backbone