PointMixer: MLP-Mixer for Point Cloud Understanding
arXiv:2111.11187
Abstract
MLP-Mixer has newly appeared as a new challenger against the realm of CNNs and transformer. Despite its simplicity compared to transformer, the concept of channel-mixing MLPs and token-mixing MLPs achieves noticeable performance in visual recognition tasks. Unlike images, point clouds are inherently sparse, unordered and irregular, which limits the direct use of MLP-Mixer for point cloud understanding. In this paper, we propose PointMixer, a universal point set operator that facilitates information sharing among unstructured 3D points. By simply replacing token-mixing MLPs with a softmax function, PointMixer can "mix" features within/between point sets. By doing so, PointMixer can be broadly used in the network as inter-set mixing, intra-set mixing, and pyramid mixing. Extensive experiments show the competitive or superior performance of PointMixer in semantic segmentation, classification, and point reconstruction against transformer-based methods.
Accepted to ECCV 2022
References in corpus (16)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- MLP-Mixer: An all-MLP Architecture for Vision
- Open3D: A Modern Library for 3D Data Processing
- Transformer in Transformer
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Submanifold Sparse Convolutional Networks
- CycleMLP: A MLP-like Architecture for Dense Prediction
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- S-MLPv2: Improved Spatial-Shift MLP Architecture for Vision
- S-MLP: Spatial-Shift MLP Architecture for Vision
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- Deep Point Cloud Reconstruction
- Hire-MLP: Vision MLP via Hierarchical Rearrangement