Efficient inference in occlusion-aware generative models of images
arXiv:1511.06362
Abstract
We present a generative model of images based on layering, in which image layers are individually generated, then composited from front to back. We are thus able to factor the appearance of an image into the appearance of individual objects within the image --- and additionally for each individual object, we can factor content from pose. Unlike prior work on layered models, we learn a shape prior for each object/layer, allowing the model to tease out which object is in front by looking for a consistent shape, without needing access to motion cues or any labeled data. We show that ordinary stochastic gradient variational bayes (SGVB), which optimizes our fully differentiable lower-bound on the log-likelihood, is sufficient to learn an interpretable representation of images. Finally we present experiments demonstrating the effectiveness of the model for inferring foreground and background objects in images.
10 pages
References in corpus (7)
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- DRAW: A Recurrent Neural Network For Image Generation
- Deep Convolutional Inverse Graphics Network
- Simultaneous Detection and Segmentation
- Discovering Hidden Factors of Variation in Deep Networks
- Spatial Transformer Networks
- Amodal Completion and Size Constancy in Natural Scenes
Cited by in corpus (13)
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
- Attend, Infer, Repeat: Fast Scene Understanding with Generative Models
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations
- Deep Successor Reinforcement Learning
- Unsupervised part representation by Flow Capsules
- GENESIS-V2: Inferring Unordered Object Representations without Iterative Refinement
- Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video
- Reconstruction Bottlenecks in Object-Centric Generative Models
- Tracking by Animation: Unsupervised Learning of Multi-Object Attentive Trackers
- Tagger: Deep Unsupervised Perceptual Grouping
- Learning Segmentation Masks with the Independence Prior
- Knowledge-Guided Object Discovery with Acquired Deep Impressions
- MarioNette: Self-Supervised Sprite Learning