Video Enhancement with Task-Oriented Flow
arXiv:1711.09078 · doi:10.1007/s11263-018-01144-2
Abstract
Many video enhancement algorithms rely on optical flow to register frames in a video sequence. Precise flow estimation is however intractable; and optical flow itself is often a sub-optimal representation for particular video processing tasks. In this paper, we propose task-oriented flow (TOFlow), a motion representation learned in a self-supervised, task-specific manner. We design a neural network with a trainable motion estimation component and a video processing component, and train them jointly to learn the task-oriented flow. For evaluation, we build Vimeo-90K, a large-scale, high-quality video dataset for low-level video processing. TOFlow outperforms traditional optical flow on standard benchmarks as well as our Vimeo-90K dataset in three video processing tasks: frame interpolation, video denoising/deblocking, and video super-resolution.
IJCV 2019. Project page: http://toflow.csail.mit.edu
References in corpus (3)
Cited by in corpus (72)
- Artificial Intelligence in the Creative Industries: A Review
- Temporal Context Mining for Learned Video Compression
- M-LVC: Multiple Frames Prediction for Learned Video Compression
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression
- Deformable 3D Convolution for Video Super-Resolution
- BVI-DVC: A Training Database for Deep Video Compression
- Scalable Image Coding for Humans and Machines
- Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model
- Deformable Non-local Network for Video Super-Resolution
- Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression
- Variational Deep Image Restoration
- Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video Denoising
- ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
- Advancing Learned Video Compression with In-loop Frame Prediction
- Towards Robust Neural Image Compression: Adversarial Attack and Model Finetuning
- Deep Learning in Physical Layer: Review on Data Driven End-to-End Communication Systems and their Enabling Semantic Applications
- Temporal Consistency Learning of inter-frames for Video Super-Resolution
- Omnidirectional Video Super-Resolution using Deep Learning
- Joint Denoising and Demosaicking with Green Channel Prior for Real-world Burst Images
- STDAN: Deformable Attention Network for Space-Time Video Super-Resolution
- Learning Cross-Scale Weighted Prediction for Efficient Neural Video Compression
- Optical Flow Reusing for High-Efficiency Space-Time Video Super Resolution
- Learned Video Compression via Heterogeneous Deformable Compensation Network
- A Survey of Deep Learning Video Super-Resolution
- iSeeBetter: Spatio-temporal video super-resolution using recurrent generative back-projection networks
- Combining Progressive Rethinking and Collaborative Learning: A Deep Framework for In-Loop Filtering
- Toward DNN of LUTs: Learning Efficient Image Restoration with Multiple Look-Up Tables
- Efficient Burst Raw Denoising with Variance Stabilization and Multi-frequency Denoising Network
- Exploring Long- and Short-Range Temporal Information for Learned Video Compression
- IBVC: Interpolation-driven B-frame Video Compression
- Neighbor Correspondence Matching for Flow-based Video Frame Synthesis
- Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
- Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video Compression
- ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video Compression
- Contextual colorization and denoising for low-light ultra high resolution sequences
- A Dual Sensor Computational Camera for High Quality Dark Videography
- BVI-AOM: A New Training Dataset for Deep Video Compression Optimization
- Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
- BVI-VFI: A Video Quality Database for Video Frame Interpolation
- TMP: Temporal Motion Propagation for Online Video Super-Resolution
- Learned Scalable Video Coding For Humans and Machines
- Learning to Compress Videos without Computing Motion
- Dynamic Frame Interpolation in Wavelet Domain
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face Videos
- Learned Wavelet Video Coding using Motion Compensated Temporal Filtering
- Real-Time Video Deblurring via Lightweight Motion Compensation
- Unlocking the Potential of Digital Pathology: Novel Baselines for Compression
- Progressive Motion Context Refine Network for Efficient Video Frame Interpolation
- Texture-aware Video Frame Interpolation
- End-to-End Optimized Image Compression with the Frequency-Oriented Transform
- Towards Real-Time Neural Video Codec for Cross-Platform Application Using Calibration Information
- RoVISQ: Reduction of Video Service Quality via Adversarial Attacks on Deep Learning-based Video Compression
- Enhanced Diagnostic Fidelity in Pathology Whole Slide Image Compression via Deep Learning
- Multi-frame Joint Enhancement for Early Interlaced Videos
- Learning-Based Conditional Image Coder Using Color Separation
- Accelerating Learnt Video Codecs with Gradient Decay and Layer-wise Distillation
- Efficient Learned Wavelet Image and Video Coding
- A New Multi-Picture Architecture for Learned Video Deinterlacing and Demosaicing with Parallel Deformable Convolution and Self-Attention Blocks
- EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution
- Diagnosing and Preventing Instabilities in Recurrent Video Processing
- IQNet: Image Quality Assessment Guided Just Noticeable Difference Prefiltering For Versatile Video Coding
- Cuboid-Net: A Multi-Branch Convolutional Neural Network for Joint Space-Time Video Super Resolution
- Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
- LVC-LGMC: Joint Local and Global Motion Compensation for Learned Video Compression
- Multi-Field De-interlacing using Deformable Convolution Residual Blocks and Self-Attention
- Multiscale Augmented Normalizing Flows for Image Compression
- SpikeCV: Open a Continuous Computer Vision Era
- Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
- Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
- Non-Uniform Exposure Imaging via Neuromorphic Shutter Control
- SATVSR: Scenario Adaptive Transformer for Cross Scenarios Video Super-Resolution
- A Large-Depth-Range Layer-Based Hologram Dataset for Machine Learning-Based 3D Computer-Generated Holography