Deformable 3D Convolution for Video Super-Resolution
arXiv:2004.02803 · doi:10.1109/LSP.2020.3013518
Abstract
The spatio-temporal information among video sequences is significant for video super-resolution (SR). However, the spatio-temporal information cannot be fully used by existing video SR methods since spatial feature extraction and temporal motion compensation are usually performed sequentially. In this paper, we propose a deformable 3D convolution network (D3Dnet) to incorporate spatio-temporal information from both spatial and temporal dimensions for video SR. Specifically, we introduce deformable 3D convolution (D3D) to integrate deformable convolution with 3D convolution, obtaining both superior spatio-temporal modeling capability and motion-aware modeling flexibility. Extensive experiments have demonstrated the effectiveness of D3D in exploiting spatio-temporal information. Comparative results show that our network achieves state-of-the-art SR performance. Code is available at: https://github.com/XinyiYing/D3Dnet.
Accepted by IEEE Signal Processing Letters
Cited by in corpus (12)
- Light Field Image Super-Resolution Using Deformable Convolution
- Local-Global Temporal Difference Learning for Satellite Video Super-Resolution
- Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
- Distortion-aware Monocular Depth Estimation for Omnidirectional Images
- Infrared Small Target Detection in Satellite Videos: A New Dataset and A Novel Recurrent Feature Refinement Framework
- STDAN: Deformable Attention Network for Space-Time Video Super-Resolution
- A Survey of Deep Learning Video Super-Resolution
- Exploiting Optical Flow Guidance for Transformer-Based Video Inpainting
- Multi-frame Joint Enhancement for Early Interlaced Videos
- A New Multi-Picture Architecture for Learned Video Deinterlacing and Demosaicing with Parallel Deformable Convolution and Self-Attention Blocks
- Symmetric Parallax Attention for Stereo Image Super-Resolution
- Facial Depth and Normal Estimation using Single Dual-Pixel Camera