Colorization as a Proxy Task for Visual Understanding
arXiv:1703.04044
Abstract
We investigate and improve self-supervision as a drop-in replacement for ImageNet pretraining, focusing on automatic colorization as the proxy task. Self-supervised training has been shown to be more promising for utilizing unlabeled data than other, traditional unsupervised learning methods. We build on this success and evaluate the ability of our self-supervised network in several contexts. On VOC segmentation and classification tasks, we present results that are state-of-the-art among methods not using ImageNet labels for pretraining representations. Moreover, we present the first in-depth analysis of self-supervision via colorization, concluding that formulation of the loss, training details and network architecture play important roles in its effectiveness. This investigation is further expanded by revisiting the ImageNet pretraining paradigm, asking questions such as: How much training data is needed? How many labels are needed? How much do features change when fine-tuned? We relate these questions back to self-supervision by showing that colorization provides a similarly powerful supervisory signal as various flavors of ImageNet pretraining.
CVPR 2017 (Project page: http://people.cs.uchicago.edu/~larsson/color-proxy/)
References in corpus (3)
Cited by in corpus (30)
- ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
- Unsupervised Representation Learning by Predicting Image Rotations
- Deep Graph Contrastive Representation Learning
- Multi-task Self-Supervised Learning for Human Activity Detection
- Unsupervised Domain Adaptation through Self-Supervision
- Improvements to context based self-supervised learning
- Test-Time Training with Self-Supervision for Generalization under Distribution Shifts
- Learning Image Representations by Completing Damaged Jigsaw Puzzles
- Leveraging Unlabeled Data for Crowd Counting by Learning to Rank
- Transitive Invariance for Self-supervised Visual Representation Learning
- SelFlow: Self-Supervised Learning of Optical Flow
- Unsupervised Feature Learning by Cross-Level Instance-Group Discrimination
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Revisiting Image Aesthetic Assessment via Self-Supervised Feature Learning
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- ClusterFit: Improving Generalization of Visual Representations
- Unsupervised Visual Attention and Invariance for Reinforcement Learning
- Are We Hungry for 3D LiDAR Data for Semantic Segmentation? A Survey and Experimental Study
- Embedding Task Knowledge into 3D Neural Networks via Self-supervised Learning
- Fully Automatic Video Colorization with Self-Regularization and Diversity
- Label-efficient audio classification through multitask learning and self-supervision
- Conditional Alignment and Uniformity for Contrastive Learning with Continuous Proxy Labels
- Robust contrastive learning and nonlinear ICA in the presence of outliers
- Learning to Look Around: Intelligently Exploring Unseen Environments for Unknown Tasks
- Temporal Interpolation as an Unsupervised Pretraining Task for Optical Flow Estimation
- ROI Regularization for Semi-supervised and Supervised Learning
- Exploit Clues from Views: Self-Supervised and Regularized Learning for Multiview Object Recognition
- Learning Rich Representations For Structured Visual Prediction Tasks