Flow-Guided Feature Aggregation for Video Object Detection
arXiv:1703.10025
Abstract
Extending state-of-the-art object detectors from image to video is challenging. The accuracy of detection suffers from degenerated object appearances in videos, e.g., motion blur, video defocus, rare poses, etc. Existing work attempts to exploit temporal information on box level, but such methods are not trained end-to-end. We present flow-guided feature aggregation, an accurate and end-to-end learning framework for video object detection. It leverages temporal coherence on feature level instead. It improves the per-frame features by aggregation of nearby features along the motion paths, and thus improves the video recognition accuracy. Our method significantly improves upon strong single-frame baselines in ImageNet VID, especially for more challenging fast moving objects. Our framework is principled, and on par with the best engineered systems winning the ImageNet VID challenges 2016, without additional bells-and-whistles. The proposed method, together with Deep Feature Flow, powered the winning entry of ImageNet VID challenges 2017. The code is available at https://github.com/msracver/Flow-Guided-Feature-Aggregation.
References in corpus (22)
- Deep Residual Learning for Image Recognition
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- R-FCN: Object Detection via Region-based Fully Convolutional Networks
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
- Going Deeper with Convolutions
- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos
- Rich feature hierarchies for accurate object detection and semantic segmentation
- Deformable Convolutional Networks
- Action Recognition using Visual Attention
- Describing Videos by Exploiting Temporal Structure
- Learning Spatiotemporal Features with 3D Convolutional Networks
- Seq-NMS for Video Object Detection
- Delving Deeper into Convolutional Networks for Learning Video Representations
- Human Action Recognition using Factorized Spatio-Temporal Convolutional Networks
- VideoLSTM Convolves, Attends and Flows for Action Recognition
- STFCN: Spatio-Temporal FCN for Semantic Video Segmentation
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- Cosine Normalization: Using Cosine Similarity Instead of Dot Product in Neural Networks
- Optical Flow Estimation using a Spatial Pyramid Network
- On The Stability of Video Detection and Tracking
- Deep End2End Voxel2Voxel Prediction
- Deep Feature Flow for Video Recognition
Cited by in corpus (34)
- Deep Learning for UAV-based Object Detection and Tracking: A Survey
- Deformable ConvNets v2: More Deformable, Better Results
- Simple Baselines for Human Pose Estimation and Tracking
- MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
- Integrated Object Detection and Tracking with Tracklet-Conditioned Detection
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- QueryProp: Object Query Propagation for High-Performance Video Object Detection
- Towards High Performance Video Object Detection for Mobiles
- Memory Enhanced Global-Local Aggregation for Video Object Detection
- Relation Distillation Networks for Video Object Detection
- Impression Network for Video Object Detection
- Transferable Adversarial Attacks for Image and Video Object Detection
- End-to-end Flow Correlation Tracking with Spatial-temporal Attention
- A semi-supervised self-training method to develop assistive intelligence for segmenting multiclass bridge elements from inspection videos
- Learning Discriminative Motion Features Through Detection
- Occlusion Aware Unsupervised Learning of Optical Flow
- Video Instance Segmentation
- Optimizing Video Object Detection via a Scale-Time Lattice
- Cross View Fusion for 3D Human Pose Estimation
- Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation
- Fast Object Detection in Compressed Video
- Towards High Performance Video Object Detection
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- Privid: Practical, Privacy-Preserving Video Analytics Queries
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Left Ventricle Segmentation via Optical-Flow-Net from Short-axis Cine MRI: Preserving the Temporal Coherence of Cardiac Motion
- Efficient Uncertainty Estimation for Semantic Segmentation in Videos
- Detect or Track: Towards Cost-Effective Video Object Detection/Tracking
- Adaptive Temporal Encoding Network for Video Instance-level Human Parsing
- Recurrent Flow-Guided Semantic Forecasting
- Real time expert system for anomaly detection of aerators based on computer vision technology and existing surveillance cameras
- Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
- Plug & Play Convolutional Regression Tracker for Video Object Detection
- Long Short-Term Relation Networks for Video Action Detection