Disentanglement of Latent Representations via Causal Interventions
arXiv:2302.00869 · doi:10.24963/ijcai.2023/361
Abstract
The process of generating data such as images is controlled by independent and unknown factors of variation. The retrieval of these variables has been studied extensively in the disentanglement, causal representation learning, and independent component analysis fields. Recently, approaches merging these domains together have shown great success. Instead of directly representing the factors of variation, the problem of disentanglement can be seen as finding the interventions on one image that yield a change to a single factor. Following this assumption, we introduce a new method for disentanglement inspired by causal dynamics that combines causality theory with vector-quantized variational autoencoders. Our model considers the quantized vectors as causal variables and links them in a causal graph. It performs causal interventions on the graph and generates atomic transitions affecting a unique factor of variation in the image. We also introduce a new task of action retrieval that consists of finding the action responsible for the transition between two images. We test our method on standard synthetic and real-world disentanglement datasets. We show that it can effectively disentangle the factors of variation and perform precise interventions on high-level semantic attributes of an image without affecting its quality, even with imbalanced data distributions.
16 pages, 10 pages for the main paper and 6 pages for the supplement, 14 figures, accepted to IJCAI 2023. V3: content matches the IJCAI version
References in corpus (7)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Zero-Shot Text-to-Image Generation
- Weakly-Supervised Disentanglement Without Compromises
- Deep End-to-end Causal Inference
- Variational Causal Networks: Approximate Bayesian Inference over Causal Structures
- Interventional Causal Representation Learning
- Discrete Key-Value Bottleneck