Learning Representation for Clustering via Prototype Scattering and Positive Sampling
arXiv:2111.11821 · doi:10.1109/TPAMI.2022.3216454
Abstract
Existing deep clustering methods rely on either contrastive or non-contrastive representation learning for downstream clustering task. Contrastive-based methods thanks to negative pairs learn uniform representations for clustering, in which negative pairs, however, may inevitably lead to the class collision issue and consequently compromise the clustering performance. Non-contrastive-based methods, on the other hand, avoid class collision issue, but the resulting non-uniform representations may cause the collapse of clustering. To enjoy the strengths of both worlds, this paper presents a novel end-to-end deep clustering method with prototype scattering and positive sampling, termed ProPos. Specifically, we first maximize the distance between prototypical representations, named prototype scattering loss, which improves the uniformity of representations. Second, we align one augmented view of instance with the sampled neighbors of another view -- assumed to be truly positive pair in the embedding space -- to improve the within-cluster compactness, termed positive sampling alignment. The strengths of ProPos are avoidable class collision issue, uniform representations, well-separated clusters, and within-cluster compactness. By optimizing ProPos in an end-to-end expectation-maximization framework, extensive experimental results demonstrate that ProPos achieves competing performance on moderate-scale clustering benchmark datasets and establishes new state-of-the-art performance on large-scale datasets. Source code is available at \url{https://github.com/Hzzone/ProPos}.
Accepted by TPAMI 2022
References in corpus (8)
- Bootstrap your own latent: A new approach to self-supervised Learning
- A Theoretical Analysis of Contrastive Unsupervised Representation Learning
- SPICE: Semantic Pseudo-labeling for Image Clustering
- Twin Contrastive Learning for Online Clustering
- AutoNovel: Automatically Discovering and Learning Novel Visual Categories
- Clustering-friendly Representation Learning via Instance Discrimination and Feature Decorrelation
- Incremental False Negative Detection for Contrastive Learning
- How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive Learning
Cited by in corpus (3)
- Reliable Representation Learning for Incomplete Multi-View Missing Multi-Label Classification
- Forget Less, Count Better: A Domain-Incremental Self-Distillation Learning Benchmark for Lifelong Crowd Counting
- EASTER: Embedding Aggregation-based Heterogeneous Models Training in Vertical Federated Learning