UIU-Net: U-Net in U-Net for Infrared Small Object Detection
arXiv:2212.00968 · doi:10.1109/TIP.2022.3228497
Abstract
Learning-based infrared small object detection methods currently rely heavily on the classification backbone network. This tends to result in tiny object loss and feature distinguishability limitations as the network depth increases. Furthermore, small objects in infrared images are frequently emerged bright and dark, posing severe demands for obtaining precise object contrast information. For this reason, we in this paper propose a simple and effective ``U-Net in U-Net'' framework, UIU-Net for short, and detect small objects in infrared images. As the name suggests, UIU-Net embeds a tiny U-Net into a larger U-Net backbone, enabling the multi-level and multi-scale representation learning of objects. Moreover, UIU-Net can be trained from scratch, and the learned features can enhance global and local contrast information effectively. More specifically, the UIU-Net model is divided into two modules: the resolution-maintenance deep supervision (RM-DS) module and the interactive-cross attention (IC-A) module. RM-DS integrates Residual U-blocks into a deep supervision network to generate deep multi-scale resolution-maintenance features while learning global context information. Further, IC-A encodes the local context information between the low-level details and high-level semantic features. Extensive experiments conducted on two infrared single-frame image datasets, i.e., SIRST and Synthetic datasets, show the effectiveness and superiority of the proposed UIU-Net in comparison with several state-of-the-art infrared small object detection methods. The proposed UIU-Net also produces powerful generalization performance for video sequence infrared small object datasets, e.g., ATR ground/air video sequence dataset. The codes of this work are available openly at \url{https://github.com/danfenghong/IEEE_TIP_UIU-Net}.
References in corpus (5)
- Fully Convolutional Networks for Semantic Segmentation
- More Diverse Means Better: Multimodal Deep Learning Meets Remote Sensing Imagery Classification
- Deep Learning for UAV-based Object Detection and Tracking: A Survey
- FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic Segmentation
- TBC-Net: A real-time detector for infrared small target detection using semantic constraint
Cited by in corpus (21)
- Multimodal Fusion Transformer for Remote Sensing Image Classification
- SCTransNet: Spatial-channel Cross Transformer Network for Infrared Small Target Detection
- Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection with Single Point Supervision
- A Comprehensive Survey for Hyperspectral Image Classification: The Evolution from Conventional to Transformers and Mamba Models
- Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection
- YOLO-MST: Multiscale deep learning method for infrared small target detection based on super-resolution and YOLO
- DATransNet: Dynamic Attention Transformer Network for Infrared Small Target Detection
- Infrared Small Target Detection in Satellite Videos: A New Dataset and A Novel Recurrent Feature Refinement Framework
- Deep learning based infrared small object segmentation: Challenges and future directions
- Unleashing the Power of Generic Segmentation Models: A Simple Baseline for Infrared Small Target Detection
- Deep-NFA: a Deep Framework for Small Object Detection
- RRCANet: Recurrent Reusable-Convolution Attention Network for Infrared Small Target Detection
- UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning
- Semi-supervised Semantic Segmentation for Remote Sensing Images via Multi-scale Uncertainty Consistency and Cross-Teacher-Student Attention
- SDS-Net: Shallow-Deep Synergism-detection Network for infrared small target detection
- Transformers Fusion across Disjoint Samples for Hyperspectral Image Classification
- Automatic Extraction of Road Networks by using Teacher-Student Adaptive Structural Deep Belief Network and Its Application to Landslide Disaster
- Weakly-supervised Contrastive Learning with Quantity Prompts for Moving Infrared Small Target Detection
- MSCA-Net:Multi-Scale Context Aggregation Network for Infrared Small Target Detection
- DCCS-Det: Directional Context and Cross-Scale-Aware Detector for Infrared Small Target
- Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review