Beyond Cats and Dogs: Semi-supervised Classification of fuzzy labels with overclustering
arXiv:2012.01768 · doi:10.3390/s21196661
Abstract
A long-standing issue with deep learning is the need for large and consistently labeled datasets. Although the current research in semi-supervised learning can decrease the required amount of annotated data by a factor of 10 or even more, this line of research still uses distinct classes like cats and dogs. However, in the real-world we often encounter problems where different experts have different opinions, thus producing fuzzy labels. We propose a novel framework for handling semi-supervised classifications of such fuzzy labels. Our framework is based on the idea of overclustering to detect substructures in these fuzzy labels. We propose a novel loss to improve the overclustering capability of our framework and show on the common image classification dataset STL-10 that it is faster and has better overclustering performance than previous work. On a real-world plankton dataset, we illustrate the benefit of overclustering for fuzzy labels and show that we beat previous state-of-the-art semisupervised methods. Moreover, we acquire 5 to 10% more consistent predictions of substructures.
Reworked version available at arXiv:2110.06630, Published in Sensors 2021 (see DOI link)
References in corpus (8)
- A Simple Framework for Contrastive Learning of Visual Representations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- Temporal Ensembling for Semi-Supervised Learning
- DivideMix: Learning with Noisy Labels as Semi-supervised Learning
- A survey on Semi-, Self- and Unsupervised Learning for Image Classification
- A Realistic Fish-Habitat Dataset to Evaluate Algorithms for Underwater Visual Analysis
- Learning from Noisy Labels with Deep Neural Networks: A Survey
- Deep learning with self-supervision and uncertainty regularization to count fish in underwater images