Self-supervised visual learning in the low-data regime: a comparative evaluation
arXiv:2404.17202 · doi:10.1016/j.neucom.2024.129199
Abstract
Self-Supervised Learning (SSL) is a valuable and robust training methodology for contemporary Deep Neural Networks (DNNs), enabling unsupervised pretraining on a 'pretext task' that does not require ground-truth labels/annotation. This allows efficient representation learning from massive amounts of unlabeled training data, which in turn leads to increased accuracy in a 'downstream task' by exploiting supervised transfer learning. Despite the relatively straightforward conceptualization and applicability of SSL, it is not always feasible to collect and/or to utilize very large pretraining datasets, especially when it comes to real-world application settings. In particular, in cases of specialized and domain-specific application scenarios, it may not be achievable or practical to assemble a relevant image pretraining dataset in the order of millions of instances or it could be computationally infeasible to pretrain at this scale, e.g., due to unavailability of sufficient computational resources that SSL methods typically require to produce improved visual analysis results. This situation motivates an investigation on the effectiveness of common SSL pretext tasks, when the pretraining dataset is of relatively limited/constrained size. This work briefly introduces the main families of modern visual SSL methods and, subsequently, conducts a thorough comparative experimental evaluation in the low-data regime, targeting to identify: a) what is learnt via low-data SSL pretraining, and b) how do different SSL categories behave in such training scenarios. Interestingly, for domain-specific downstream tasks, in-domain low-data SSL pretraining outperforms the common approach of large-scale pretraining on general datasets.
Article published in Elsevier's Neurocomputing journal: https://www.sciencedirect.com/science/article/pii/S0925231224019702
References in corpus (12)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- DINOv2: Learning Robust Visual Features without Supervision
- BEiT: BERT Pre-Training of Image Transformers
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- HIT-UAV: A high-altitude infrared thermal dataset for Unmanned Aerial Vehicle-based object detection
- iBOT: Image BERT Pre-Training with Online Tokenizer
- Self-Supervised Learning for Videos: A Survey
- Self-supervised Pretraining of Visual Features in the Wild
- Leveraging Single-View Images for Unsupervised 3D Point Cloud Completion
- Deep Clustering with Features from Self-Supervised Pretraining
- Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization