Learning Multi-Domain Convolutional Neural Networks for Visual Tracking
arXiv:1510.07945
Abstract
We propose a novel visual tracking algorithm based on the representations from a discriminatively trained Convolutional Neural Network (CNN). Our algorithm pretrains a CNN using a large set of videos with tracking ground-truths to obtain a generic target representation. Our network is composed of shared layers and multiple branches of domain-specific layers, where domains correspond to individual training sequences and each branch is responsible for binary classification to identify the target in each domain. We train the network with respect to each domain iteratively to obtain generic target representations in the shared layers. When tracking a target in a new sequence, we construct a new network by combining the shared layers in the pretrained CNN with a new binary classification layer, which is updated online. Online tracking is performed by evaluating the candidate windows randomly sampled around the previous target state. The proposed algorithm illustrates outstanding performance compared with state-of-the-art methods in existing tracking benchmarks.
References in corpus (5)
Cited by in corpus (30)
- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos
- Deep Reinforcement Learning for Visual Object Tracking in Videos
- Learning Background-Aware Correlation Filters for Visual Tracking
- A Twofold Siamese Network for Real-Time Object Tracking
- Learning Video Object Segmentation from Static Images
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking
- Spatially Supervised Recurrent Convolutional Neural Networks for Visual Object Tracking
- Siamese Cascaded Region Proposal Networks for Real-Time Visual Tracking
- Describe and Attend to Track: Learning Natural Language guided Structural Representation and Visual Attention for Object Tracking
- Fully-Convolutional Siamese Networks for Object Tracking
- Correlated and Individual Multi-Modal Deep Learning for RGB-D Object Recognition
- CompRRAE: RRAM-based Convolutional Neural Network Accelerator with Reduced Computations through a Runtime Activation Estimation
- Euphrates: Algorithm-SoC Co-Design for Low-Power Mobile Continuous Vision
- Recurrent Filter Learning for Visual Tracking
- Real-time visual tracking by deep reinforced decision making
- CyLKs: Unsupervised Cycle Lucas-Kanade Network for Landmark Tracking
- Towards a Better Match in Siamese Network Based Visual Object Tracker
- Joint Representation and Truncated Inference Learning for Correlation Filter based Tracking
- Fast and Accurate Online Video Object Segmentation via Tracking Parts
- Visual Tracking via Dynamic Graph Learning
- Learning to Select Pre-Trained Deep Representations with Bayesian Evidence Framework
- Human Action Adverb Recognition: ADHA Dataset and A Three-Stream Hybrid Model
- Deep-LK for Efficient Adaptive Object Tracking
- Feature Selection Convolutional Neural Networks for Visual Tracking
- Visual Data Augmentation through Learning
- Learning Compact Target-Oriented Feature Representations for Visual Tracking
- Deep Learning based Multi-Modal Sensing for Tracking and State Extraction of Small Quadcopters
- Active Collaborative Ensemble Tracking
- Revisiting the details when evaluating a visual tracker
- Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex Manifold