Unsupervised Learning of Spatiotemporally Coherent Metrics
arXiv:1412.6056
Abstract
Current state-of-the-art classification and detection algorithms rely on supervised training. In this work we study unsupervised feature learning in the context of temporally coherent video data. We focus on feature learning from unlabeled video data, using the assumption that adjacent video frames contain semantically similar information. This assumption is exploited to train a convolutional pooling auto-encoder regularized by slowness and sparsity. We establish a connection between slow feature learning to metric learning and show that the trained encoder can be used to define a more temporally and semantically coherent metric.
To appear at ICCV2015
References in corpus (1)
Cited by in corpus (9)
- Unsupervised Learning of Depth and Ego-Motion from Video
- Unsupervised Learning of Visual Representations using Videos
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- 3DMatch: Learning Local Geometric Descriptors from RGB-D Reconstructions
- Temporal Generative Adversarial Nets with Singular Value Clipping
- Unsupervised Learning for Large-Scale Fiber Detection and Tracking in Microscopic Material Images
- Learning image representations tied to ego-motion
- Competitive Learning Enriches Learning Representation and Accelerates the Fine-tuning of CNNs
- Biologically-Motivated Deep Learning Method using Hierarchical Competitive Learning