SwinNet: Swin Transformer drives edge-aware RGB-D and RGB-T salient object detection
arXiv:2204.05585 · doi:10.1109/TCSVT.2021.3127149
Abstract
Convolutional neural networks (CNNs) are good at extracting contexture features within certain receptive fields, while transformers can model the global long-range dependency features. By absorbing the advantage of transformer and the merit of CNN, Swin Transformer shows strong feature representation ability. Based on it, we propose a cross-modality fusion model SwinNet for RGB-D and RGB-T salient object detection. It is driven by Swin Transformer to extract the hierarchical features, boosted by attention mechanism to bridge the gap between two modalities, and guided by edge information to sharp the contour of salient object. To be specific, two-stream Swin Transformer encoder first extracts multi-modality features, and then spatial alignment and channel re-calibration module is presented to optimize intra-level cross-modality features. To clarify the fuzzy boundary, edge-guided decoder achieves inter-level cross-modality fusion under the guidance of edge features. The proposed model outperforms the state-of-the-art models on RGB-D and RGB-T datasets, showing that it provides more insight into the cross-modality complementarity task.
Online published in TCSVT
References in corpus (10)
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Online Tracking by Learning Discriminative Saliency Map with Convolutional Neural Network
- Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing Images
- Multi-interactive Dual-decoder for RGB-thermal Salient Object Detection
- A Parallel Down-Up Fusion Network for Salient Object Detection in Optical Remote Sensing Images
- Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object Detection
- RGBT Salient Object Detection: A Large-scale Dataset and Benchmark
- DUT-LFSaliency: Versatile Dataset and Light Field-to-RGB Saliency Detection
- CAT: Cross Attention in Vision Transformer
- CMA-Net: A Cascaded Mutual Attention Network for Light Field Salient Object Detection
Cited by in corpus (13)
- HRTransNet: HRFormer-Driven Two-Modality Salient Object Detection
- Salient Object Detection in Optical Remote Sensing Images Driven by Transformer
- Position-Aware Relation Learning for RGB-Thermal Salient Object Detection
- Swin-transformer-yolov5 For Real-time Wine Grape Bunch Detection
- VST++: Efficient and Stronger Visual Saliency Transformer
- Quality-aware Selective Fusion Network for V-D-T Salient Object Detection
- FDiff-Fusion:Denoising diffusion fusion network based on fuzzy learning for 3D medical image segmentation
- METER: a mobile vision transformer architecture for monocular depth estimation
- CST-YOLO: A Novel Method for Blood Cell Detection Based on Improved YOLOv7 and CNN-Swin Transformer
- Spectrum-driven Mixed-frequency Network for Hyperspectral Salient Object Detection
- UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning
- Deep Fourier-embedded Network for RGB and Thermal Salient Object Detection
- Efficient Fourier Filtering Network with Contrastive Learning for AAV-based Unaligned Bimodal Salient Object Detection