Boxhead: A Dataset for Learning Hierarchical Representations
arXiv:2110.03628
Abstract
Disentanglement is hypothesized to be beneficial towards a number of downstream tasks. However, a common assumption in learning disentangled representations is that the data generative factors are statistically independent. As current methods are almost solely evaluated on toy datasets where this ideal assumption holds, we investigate their performance in hierarchical settings, a relevant feature of real-world data. In this work, we introduce Boxhead, a dataset with hierarchically structured ground-truth generative factors. We use this novel dataset to evaluate the performance of state-of-the-art autoencoder-based disentanglement models and observe that hierarchical models generally outperform single-layer VAEs in terms of disentanglement of hierarchically arranged factors.
NeurIPS 2021 Workshop on Shared Visual Representations in Human and Machine Intelligence (SVRHM 2021)
References in corpus (11)
- Open3D: A Modern Library for 3D Data Processing
- NVAE: A Deep Hierarchical Variational Autoencoder
- Flexibly Fair Representation Learning by Disentanglement
- Weakly-Supervised Disentanglement Without Compromises
- Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
- Unsupervised Discovery of Parts, Structure, and Dynamics
- Disentangled State Space Representations
- Progressive Learning and Disentanglement of Hierarchical Representations
- S2RMs: Spatially Structured Recurrent Modules
- Benchmarks, Algorithms, and Metrics for Hierarchical Disentanglement
- Generalization and Robustness Implications in Object-Centric Learning