Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects
arXiv:1806.01794
Abstract
We present Sequential Attend, Infer, Repeat (SQAIR), an interpretable deep generative model for videos of moving objects. It can reliably discover and track objects throughout the sequence of frames, and can also generate future frames conditioning on the current frame, thereby simulating expected motion of objects. This is achieved by explicitly encoding object presence, locations and appearances in the latent variables of the model. SQAIR retains all strengths of its predecessor, Attend, Infer, Repeat (AIR, Eslami et. al., 2016), including learning in an unsupervised manner, and addresses its shortcomings. We use a moving multi-MNIST dataset to show limitations of AIR in detecting overlapping or partially occluded objects, and show how SQAIR overcomes them by leveraging temporal consistency of objects. Finally, we also apply SQAIR to real-world pedestrian CCTV data, where it learns to reliably detect, track and generate walking pedestrians with no supervision.
25 pages, 19 figures, NeurIPS 2018, code: https://github.com/akosiorek/sqair, video: https://youtu.be/-IUNQgSLE0c
Cited by in corpus (53)
- Object-Centric Learning with Slot Attention
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations
- BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
- Contrastive Learning of Structured World Models
- Is Attention Better Than Matrix Decomposition?
- On the Binding Problem in Artificial Neural Networks
- Learning Physical Graph Representations from Visual Scenes
- RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent Spaces
- Stacked Capsule Autoencoders
- From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence
- Investigating Object Compositionality in Generative Adversarial Networks
- Entity Abstraction in Visual Model-Based Reinforcement Learning
- Decomposing 3D Scenes into Objects via Unsupervised Volume Segmentation
- Towards causal generative scene models via competition of experts
- Unsupervised Discovery of Object Radiance Fields
- Structured Object-Aware Physics Prediction for Video Modeling and Planning
- GENESIS-V2: Inferring Unordered Object Representations without Iterative Refinement
- A Perspective on Objects and Systematic Generalization in Model-Based RL
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Unsupervised Discovery of 3D Physical Objects from Video
- PDE-Driven Spatiotemporal Disentanglement
- Unsupervised Learning of Compositional Energy Concepts
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
- Generative Neurosymbolic Machines
- Improving Generative Imagination in Object-Centric World Models
- Forward Prediction for Physical Reasoning
- Occlusion resistant learning of intuitive physics from videos
- Reconstruction Bottlenecks in Object-Centric Generative Models
- Multi-objects Generation with Amortized Structural Regularization
- Learning Object-Centric Video Models by Contrasting Sets
- Towards Unsupervised Learning of Generative Models for 3D Controllable Image Synthesis
- End-to-end Recurrent Multi-Object Tracking and Trajectory Prediction with Relational Reasoning
- Slot Contrastive Networks: A Contrastive Approach for Representing Objects
- Generalization and Robustness Implications in Object-Centric Learning
- Benchmarking Unsupervised Object Representations for Video Sequences
- Unsupervised Part Discovery from Contrastive Reconstruction
- Uncovering Closed-form Governing Equations of Nonlinear Dynamics from Videos
- Unsupervised Object Keypoint Learning using Local Spatial Predictability
- Self-supervised Visual Reinforcement Learning with Object-centric Representations
- Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders
- Learning Disentangled Representations of Video with Missing Data
- Learning to Manipulate Individual Objects in an Image
- Language-Mediated, Object-Centric Representation Learning
- AutoTrajectory: Label-free Trajectory Extraction and Prediction from Videos using Dynamic Points
- Knowledge-Guided Object Discovery with Acquired Deep Impressions
- Mind the Gap when Conditioning Amortised Inference in Sequential Latent-Variable Models
- Self-Supervised Decomposition, Disentanglement and Prediction of Video Sequences while Interpreting Dynamics: A Koopman Perspective
- Unsupervised Object Learning via Common Fate
- Deep Variational Luenberger-type Observer for Stochastic Video Prediction
- Opening up Open-World Tracking
- Illiterate DALL-E Learns to Compose
- Conditional Object-Centric Learning from Video