Multimodal Fusion Transformer for Remote Sensing Image Classification
arXiv:2203.16952 · doi:10.1109/TGRS.2023.3286826
Abstract
Vision transformers (ViTs) have been trending in image classification tasks due to their promising performance when compared to convolutional neural networks (CNNs). As a result, many researchers have tried to incorporate ViTs in hyperspectral image (HSI) classification tasks. To achieve satisfactory performance, close to that of CNNs, transformers need fewer parameters. ViTs and other similar transformers use an external classification (CLS) token which is randomly initialized and often fails to generalize well, whereas other sources of multimodal datasets, such as light detection and ranging (LiDAR) offer the potential to improve these models by means of a CLS. In this paper, we introduce a new multimodal fusion transformer (MFT) network which comprises a multihead cross patch attention (mCrossPA) for HSI land-cover classification. Our mCrossPA utilizes other sources of complementary information in addition to the HSI in the transformer encoder to achieve better generalization. The concept of tokenization is used to generate CLS and HSI patch tokens, helping to learn a {distinctive representation} in a reduced and hierarchical feature space. Extensive experiments are carried out on {widely used benchmark} datasets {i.e.,} the University of Houston, Trento, University of Southern Mississippi Gulfpark (MUUFL), and Augsburg. We compare the results of the proposed MFT model with other state-of-the-art transformers, classical CNNs, and conventional classifiers models. The superior performance achieved by the proposed model is due to the use of multihead cross patch attention. The source code will be made available publicly at \url{https://github.com/AnkurDeria/MFT}.}
Published in IEEE Transactions on Geoscience and Remote Sensing
References in corpus (4)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- More Diverse Means Better: Multimodal Deep Learning Meets Remote Sensing Imagery Classification
- UIU-Net: U-Net in U-Net for Infrared Small Object Detection
- Classification of Hyperspectral and LiDAR Data Using Coupled CNNs
Cited by in corpus (15)
- Deep Hyperspectral Unmixing using Transformer Network
- Robust Multimodal Learning with Missing Modalities via Parameter-Efficient Adaptation
- FedFusion: Manifold Driven Federated Learning for Multi-satellite and Multi-modality Fusion
- A Data-Driven Review of Remote Sensing-Based Data Fusion in Precision Agriculture from Foundational to Transformer-Based Techniques
- Fusion of Satellite Images and Weather Data with Transformer Networks for Downy Mildew Disease Detection
- PolSAR Image Classification using a Hybrid Complex-Valued Network (HybridCVNet)
- On Advances, Challenges and Potentials of Remote Sensing Image Analysis in Marine Debris and Suspected Plastics Monitoring
- Spatial Gated Multi-Layer Perceptron for Land Use and Land Cover Mapping
- Enhanced Astronomical Source Classification with Integration of Attention Mechanisms and Vision Transformers
- A CNN with Noise Inclined Module and Denoise Framework for Hyperspectral Image Classification
- BihoT: A Large-Scale Dataset and Benchmark for Hyperspectral Camouflaged Object Tracking
- HyperPointFormer: Multimodal Fusion in 3D Space with Dual-Branch Cross-Attention Transformers
- MixerSENet: A Lightweight Framework for Efficient Hyperspectral Image Classification
- Graph Transformer with Disease Subgraph Positional Encoding for Improved Comorbidity Prediction
- Hyperspectral Image Classification using Spectral-Spatial Mixer Network