Deep Reinforcement Learning for Visual Object Tracking in Videos
arXiv:1701.08936
Abstract
In this paper we introduce a fully end-to-end approach for visual tracking in videos that learns to predict the bounding box locations of a target object at every frame. An important insight is that the tracking problem can be considered as a sequential decision-making process and historical semantics encode highly relevant information for future decisions. Based on this intuition, we formulate our model as a recurrent convolutional neural network agent that interacts with a video overtime, and our model can be trained with reinforcement learning (RL) algorithms to learn good tracking policies that pay attention to continuous, inter-frame correlation and maximize tracking performance in the long run. The proposed tracking algorithm achieves state-of-the-art performance in an existing tracking benchmark and operates at frame-rates faster than real-time. To the best of our knowledge, our tracker is the first neural-network tracker that combines convolutional and recurrent networks with RL algorithms.
References in corpus (18)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Playing Atari with Deep Reinforcement Learning
- DeepPose: Human Pose Estimation via Deep Neural Networks
- Perceptual Losses for Real-Time Style Transfer and Super-Resolution
- Recurrent Models of Visual Attention
- Fully Convolutional Networks for Semantic Segmentation
- Multiple Object Recognition with Visual Attention
- Rich feature hierarchies for accurate object detection and semantic segmentation
- Transferring Rich Feature Hierarchies for Robust Visual Tracking
- Learning Multi-Domain Convolutional Neural Networks for Visual Tracking
- Understanding and Diagnosing Visual Tracking Systems
- First Step toward Model-Free, Anonymous Object Tracking with Recurrent Neural Networks
- On Learning Where To Look
- Spatially Supervised Recurrent Convolutional Neural Networks for Visual Object Tracking
- Robust Visual Tracking via Convolutional Networks
Cited by in corpus (18)
- Deep Learning for Visual Tracking: A Comprehensive Survey
- Integrated Object Detection and Tracking with Tracklet-Conditioned Detection
- Learning Policies for Adaptive Tracking with Deep Feature Cascades
- S3D: Single Shot multi-Span Detector via Fully 3D Convolutional Networks
- Natural Environment Benchmarks for Reinforcement Learning
- Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-based Beam Search
- iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Learning to Track Dynamic Targets in Partially Known Environments
- DAWN: Dual Augmented Memory Network for Unsupervised Video Object Tracking
- Robust and customized methods for real-time hand gesture recognition under object-occlusion
- Deep Reinforcement Learning for Active Human Pose Estimation
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- Long and Short Memory Balancing in Visual Co-Tracking using Q-Learning
- Handcrafted and Deep Trackers: Recent Visual Object Tracking Approaches and Trends
- Weakly Supervised Video Summarization by Hierarchical Reinforcement Learning
- Real-time tracker with fast recovery from target loss
- Hide and Seek tracker: Real-time recovery from target loss