Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction
arXiv:1808.00758 · doi:10.1007/s11263-019-01217-w
Abstract
We study the problem of recovering an underlying 3D shape from a set of images. Existing learning based approaches usually resort to recurrent neural nets, e.g., GRU, or intuitive pooling operations, e.g., max/mean poolings, to fuse multiple deep features encoded from input images. However, GRU based approaches are unable to consistently estimate 3D shapes given different permutations of the same set of input images as the recurrent unit is permutation variant. It is also unlikely to refine the 3D shape given more images due to the long-term memory loss of GRU. Commonly used pooling approaches are limited to capturing partial information, e.g., max/mean values, ignoring other valuable features. In this paper, we present a new feed-forward neural module, named AttSets, together with a dedicated training algorithm, named FASet, to attentively aggregate an arbitrarily sized deep feature set for multi-view 3D reconstruction. The AttSets module is permutation invariant, computationally efficient and flexible to implement, while the FASet algorithm enables the AttSets based network to be remarkably robust and generalize to an arbitrary number of input images. We thoroughly evaluate FASet and the properties of AttSets on multiple large public datasets. Extensive experiments show that AttSets together with FASet algorithm significantly outperforms existing aggregation approaches.
IJCV 2019. Code and data are available at https://github.com/Yang7879/AttSets
References in corpus (4)
- Attentional Pooling for Action Recognition
- Attention-based Pyramid Aggregation Network for Visual Place Recognition
- Monocular Dense 3D Reconstruction of a Complex Dynamic Scene from Two Perspective Frames
- "Maximizing rigidity" revisited: a convex programming approach for generic 3D shape reconstruction from multiple perspective views
Cited by in corpus (22)
- Learning Semantic Segmentation of Large-Scale Point Clouds with Random Sampling
- Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images
- Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks
- Gaussian Radar Transformer for Semantic Segmentation in Noisy Radar Data
- GARNet: Global-Aware Multi-View 3D Reconstruction Network and the Cost-Performance Tradeoff
- SoftPool++: An Encoder-Decoder Network for Point Cloud Completion
- Learning Pose-invariant 3D Object Reconstruction from Single-view Images
- Radar Velocity Transformer: Single-scan Moving Object Segmentation in Noisy Radar Point Clouds
- Dense Voxel 3D Reconstruction Using a Monocular Event Camera
- Machine Learning for Detection of 3D Features using sparse X-ray data
- 3D-RETR: End-to-End Single and Multi-View 3D Reconstruction with Transformers
- DLA-Net: Learning Dual Local Attention Features for Semantic Segmentation of Large-Scale Building Facade Point Clouds
- Set-to-Sequence Methods in Machine Learning: a Review
- PointLoc: Deep Pose Regressor for LiDAR Point Cloud Localization
- RMS-FlowNet++: Efficient and Robust Multi-Scale Scene Flow Estimation for Large-Scale Point Clouds
- Refine3DNet: Scaling Precision in 3D Object Reconstruction from Multi-View RGB Images using Attention
- You Never Cluster Alone
- Learning to Reconstruct and Segment 3D Objects
- Compact Deep Aggregation for Set Retrieval
- LegoFormer: Transformers for Block-by-Block Multi-view 3D Reconstruction
- Representing Unordered Data Using Complex-Weighted Multiset Automata
- Multi-view 3D Reconstruction with Transformer