Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
arXiv:2001.03569 · doi:10.1109/TIP.2020.3016485
Abstract
Video coding, which targets to compress and reconstruct the whole frame, and feature compression, which only preserves and transmits the most critical information, stand at two ends of the scale. That is, one is with compactness and efficiency to serve for machine vision, and the other is with full fidelity, bowing to human perception. The recent endeavors in imminent trends of video compression, e.g. deep learning based coding tools and end-to-end image/video coding, and MPEG-7 compact feature descriptor standards, i.e. Compact Descriptors for Visual Search and Compact Descriptors for Video Analysis, promote the sustainable and fast development in their own directions, respectively. In this paper, thanks to booming AI technology, e.g. prediction and generation models, we carry out exploration in the new area, Video Coding for Machines (VCM), arising from the emerging MPEG standardization efforts1. Towards collaborative compression and intelligent analytics, VCM attempts to bridge the gap between feature coding for machine vision and video coding for human vision. Aligning with the rising Analyze then Compress instance Digital Retina, the definition, formulation, and paradigm of VCM are given first. Meanwhile, we systematically review state-of-the-art techniques in video compression and feature compression from the unique perspective of MPEG standardization, which provides the academic and industrial evidence to realize the collaborative compression of video and feature streams in a broad range of AI applications. Finally, we come up with potential VCM solutions, and the preliminary results have demonstrated the performance and efficiency gains. Further direction is discussed as well.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- YOLOv3: An Incremental Improvement
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Variational image compression with a scale hyperprior
- Generating Videos with Scene Dynamics
- Decomposing Motion and Content for Natural Video Sequence Prediction
- Stochastic Video Generation with a Learned Prior
- Learning What and Where to Draw
- Stochastic Adversarial Video Prediction
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Hierarchical Long-term Video Prediction without Supervision
Cited by in corpus (21)
- Scalable Image Coding for Humans and Machines
- Video Coding for Machines with Feature-Based Rate-Distortion Optimization
- Recent Standard Development Activities on Video Coding for Machines
- FitVid: Overfitting in Pixel-Level Video Prediction
- Boosting Neural Image Compression for Machines Using Latent Space Masking
- Collaborative Intelligence: Challenges and Opportunities
- Video Coding for Machine: Compact Visual Representation Compression for Intelligent Collaborative Analytics
- Learned Scalable Video Coding For Humans and Machines
- Human-Machine Collaborative Video Coding Through Cuboidal Partitioning
- A Rate-Distortion-Classification Approach for Lossy Image Compression
- Bridging the gap between image coding for machines and humans
- Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based Approach
- Machine Perception-Driven Image Compression: A Layered Generative Approach
- Learning End-to-End Lossy Image Compression: A Benchmark
- Pruned Lightweight Encoders for Computer Vision
- A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric Estimation
- Deep Video Codec Control for Vision Models
- End-to-end Compression Towards Machine Vision: Network Architecture Design and Optimization
- Analysis of Latent-Space Motion for Collaborative Intelligence
- Latent-Space Inpainting for Packet Loss Concealment in Collaborative Object Detection
- Revisit Visual Representation in Analytics Taxonomy: A Compression Perspective