Uformer: A General U-Shaped Transformer for Image Restoration
arXiv:2106.03106
Abstract
In this paper, we present Uformer, an effective and efficient Transformer-based architecture for image restoration, in which we build a hierarchical encoder-decoder network using the Transformer block. In Uformer, there are two core designs. First, we introduce a novel locally-enhanced window (LeWin) Transformer block, which performs nonoverlapping window-based self-attention instead of global self-attention. It significantly reduces the computational complexity on high resolution feature map while capturing local context. Second, we propose a learnable multi-scale restoration modulator in the form of a multi-scale spatial bias to adjust features in multiple layers of the Uformer decoder. Our modulator demonstrates superior capability for restoring details for various image restoration tasks while introducing marginal extra parameters and computational cost. Powered by these two designs, Uformer enjoys a high capability for capturing both local and global dependencies for image restoration. To evaluate our approach, extensive experiments are conducted on several image restoration tasks, including image denoising, motion deblurring, defocus deblurring and deraining. Without bells and whistles, our Uformer achieves superior or comparable performance compared with the state-of-the-art algorithms. The code and models are available at https://github.com/ZhendongWang6/Uformer.
17 pages, 13 figures
References in corpus (9)
- LocalViT: Analyzing Locality in Vision Transformers
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
- Residual Non-local Attention Networks for Image Restoration
- CvT: Introducing Convolutions to Vision Transformers
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Single Image Defocus Deblurring Using Kernel-Sharing Parallel Atrous Convolutions
Cited by in corpus (8)
- A broadband hyperspectral image sensor with high spatio-temporal resolution
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- Light Field Image Super-Resolution with Transformers
- MISSFormer: An Effective Medical Image Segmentation Transformer
- SwinIR: Image Restoration Using Swin Transformer
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction
- Eformer: Edge Enhancement based Transformer for Medical Image Denoising
- High Dynamic Range Image Reconstruction via Deep Explicit Polynomial Curve Estimation