PointSmile: Point Self-supervised Learning via Curriculum Mutual Information
arXiv:2301.12744 · doi:10.1007/s11432-023-4085-9
Abstract
Self-supervised learning is attracting wide attention in point cloud processing. However, it is still not well-solved to gain discriminative and transferable features of point clouds for efficient training on downstream tasks, due to their natural sparsity and irregularity. We propose PointSmile, a reconstruction-free self-supervised learning paradigm by maximizing curriculum mutual information (CMI) across the replicas of point cloud objects. From the perspective of how-and-what-to-learn, PointSmile is designed to imitate human curriculum learning, i.e., starting with an easy curriculum and gradually increasing the difficulty of that curriculum. To solve "how-to-learn", we introduce curriculum data augmentation (CDA) of point clouds. CDA encourages PointSmile to learn from easy samples to hard ones, such that the latent space can be dynamically affected to create better embeddings. To solve "what-to-learn", we propose to maximize both feature- and class-wise CMI, for better extracting discriminative features of point clouds. Unlike most of existing methods, PointSmile does not require a pretext task, nor does it require cross-modal data to yield rich latent representations. We demonstrate the effectiveness and robustness of PointSmile in downstream tasks including object classification and segmentation. Extensive results show that our PointSmile outperforms existing self-supervised methods, and compares favorably with popular fully-supervised methods on various standard architectures.
References in corpus (10)
- Bootstrap your own latent: A new approach to self-supervised Learning
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- PCT: Point cloud transformer
- Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP Framework
- Self-Contrastive Learning with Hard Negative Sampling for Self-supervised Point Cloud Learning
- CSDN: Cross-modal Shape-transfer Dual-refinement Network for Point Cloud Completion
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud Understanding
- Unsupervised Point Cloud Pre-Training via Occlusion Completion
- Distillation with Contrast is All You Need for Self-Supervised Point Cloud Representation Learning