Attend, Infer, Repeat: Fast Scene Understanding with Generative Models
arXiv:1603.08575
Abstract
We present a framework for efficient inference in structured image models that explicitly reason about objects. We achieve this by performing probabilistic inference using a recurrent neural network that attends to scene elements and processes them one at a time. Crucially, the model itself learns to choose the appropriate number of inference steps. We use this scheme to learn to perform inference in partially specified 2D models (variable-sized variational auto-encoders) and fully specified 3D models (probabilistic renderers). We show that such models learn to identify multiple objects - counting, locating and classifying the elements of a scene - without any supervision, e.g., decomposing 3D images with various numbers of objects in a single forward pass of a neural network. We further show that the networks produce accurate inferences when compared to supervised counterparts, and that their structure leads to improved generalization.
References in corpus (7)
- Adam: A Method for Stochastic Optimization
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Recurrent Models of Visual Attention
- DRAW: A Recurrent Neural Network For Image Generation
- Adaptive Computation Time for Recurrent Neural Networks
- The Informed Sampler: A Discriminative Approach to Bayesian Inference in Generative Computer Vision Models
- Efficient inference in occlusion-aware generative models of images
Cited by in corpus (70)
- An Introduction to Variational Autoencoders
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
- Deep Reinforcement Learning: An Overview
- Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders
- Adaptive Computation Time for Recurrent Neural Networks
- Recent Advances in Convolutional Neural Networks
- Unsupervised Learning for Physical Interaction through Video Prediction
- Multi-Object Representation Learning with Iterative Variational Inference
- Learning What and Where to Draw
- Relational Neural Expectation Maximization: Unsupervised Discovery of Objects and their Interactions
- BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images
- Learning to Decompose and Disentangle Representations for Video Prediction
- Unsupervised Learning of 3D Structure from Images
- Deep Successor Reinforcement Learning
- Towards an integration of deep learning and neuroscience
- One-Shot Generalization in Deep Generative Models
- Contrastive Learning of Structured World Models
- Adversarial Images for Variational Autoencoders
- Small Sample Learning in Big Data Era
- Normalizing Flows on Riemannian Manifolds
- Object Discovery with a Copy-Pasting GAN
- Tackling Over-pruning in Variational Autoencoders
- Language as a Latent Variable: Discrete Generative Models for Sentence Compression
- Generative Temporal Models with Memory
- A Whole Brain Probabilistic Generative Model: Toward Realizing Cognitive Architectures for Developmental Robots
- Investigating Object Compositionality in Generative Adversarial Networks
- Discrete and continuous representations and processing in deep learning: Looking forward
- Revisiting Reweighted Wake-Sleep for Models with Stochastic Control Flow
- Generalization in anti-causal learning
- R-SQAIR: Relational Sequential Attend, Infer, Repeat
- Unsupervised Discovery of Object Radiance Fields
- Gaussian Mixture Variational Autoencoder with Contrastive Learning for Multi-Label Classification
- Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
- Neural Density Estimation and Likelihood-free Inference
- Learning and Refining of Privileged Information-based RNNs for Action Recognition from Depth Sequences
- Information-theoretic Model Identification and Policy Search using Physics Engines with Application to Robotic Manipulation
- Gaussian mixture models with Wasserstein distance
- Structured Generative Models for Scene Understanding
- Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
- Object-Centric Image Generation with Factored Depths, Locations, and Appearances
- Inducing Interpretable Representations with Variational Autoencoders
- Probabilistic Programming with Programmable Variational Inference
- Learning to Zoom: a Saliency-Based Sampling Layer for Neural Networks
- Cerberus: A Multi-headed Derenderer
- Direct Optimization through for Discrete Variational Auto-Encoder
- Signal-based Bayesian Seismic Monitoring
- Generating new concepts with hybrid neuro-symbolic models
- Vid2Param: Modelling of Dynamics Parameters from Video
- Learning Direct Optimization for Scene Understanding
- Building Machines that Learn and Think for Themselves: Commentary on Lake et al., Behavioral and Brain Sciences, 2017
- Tagger: Deep Unsupervised Perceptual Grouping
- Recursive Neural Programs: Variational Learning of Image Grammars and Part-Whole Hierarchies
- Unsupervised and interpretable scene discovery with Discrete-Attend-Infer-Repeat
- Visually Grounded Compound PCFGs
- Decomposing Normal and Abnormal Features of Medical Images into Discrete Latent Codes for Content-Based Image Retrieval
- Disentangling 3D Prototypical Networks For Few-Shot Concept Learning
- Learning Disentangled Representations of Video with Missing Data
- Causal World Models by Unsupervised Deconfounding of Physical Dynamics
- Learning Segmentation Masks with the Independence Prior
- Substitute Teacher Networks: Learning with Almost No Supervision
- End-to-end Speech Recognition with Adaptive Computation Steps
- GMAIR: Unsupervised Object Detection Based on Spatial Attention and Gaussian Mixture
- Recurrent Existence Determination Through Policy Optimization
- Recurrent Attention Models with Object-centric Capsule Representation for Multi-object Recognition
- An Interpretable Generative Model for Handwritten Digit Image Synthesis
- Understanding in Artificial Intelligence
- A backward pass through a CNN using a generative model of its activations
- Disassembling Object Representations without Labels
- NP-DRAW: A Non-Parametric Structured Latent Variable Model for Image Generation