Do Different Tracking Tasks Require Different Appearance Models?
arXiv:2107.02156
Abstract
Tracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature has fragmented too, and now novel approaches proposed by the community are usually specialised to fit only one specific setup. To understand to what extent this specialisation is necessary, in this work we present UniTrack, a solution to address five different tasks within the same framework. UniTrack consists of a single and task-agnostic appearance model, which can be learned in a supervised or self-supervised fashion, and multiple ``heads'' that address individual tasks and do not require training. We show how most tracking tasks can be solved within this framework, and that the same appearance model can be successfully used to obtain results that are competitive against specialised methods for most of the tasks considered. The framework also allows us to analyse appearance models obtained with the most recent self-supervised methods, thus extending their evaluation and comparison to a larger variety of important problems.
To appear at NeurIPS 2021
References in corpus (17)
- High-Speed Tracking with Kernelized Correlation Filters
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking
- Do Convnets Learn Correspondence?
- Self-EMD: Self-Supervised Object Detection without ImageNet
- Joint-task Self-supervised Learning for Temporal Correspondence
- Unsupervised Deep Tracking
- FEELVOS: Fast End-to-End Embedding Learning for Video Object Segmentation
- SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation
- MAT: Motion-Aware Multi-Object Tracking
- Track to Detect and Segment: An Online Multi-Object Tracker
- TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training Model
- Local Metrics for Multi-Object Tracking
- Unsupervised Deep Representation Learning for Real-Time Tracking
- Learning to Track Instances without Video Annotations