Learning Physical Intuition of Block Towers by Example
arXiv:1603.01312
Abstract
Wooden blocks are a common toy for infants, allowing them to develop motor skills and gain intuition about the physical behavior of the world. In this paper, we explore the ability of deep feed-forward models to learn such intuitive physics. Using a 3D game engine, we create small towers of wooden blocks whose stability is randomized and render them collapsing (or remaining upright). This data allows us to train large convolutional network models which can accurately predict the outcome, as well as estimating the block trajectories. The models are also able to generalize in two important ways: (i) to new physical scenarios, e.g. towers with an additional block and (ii) to images of real wooden blocks, where it obtains a performance comparable to human subjects.
References in corpus (6)
- Continuous control with deep reinforcement learning
- Going Deeper with Convolutions
- Learning to Segment Object Candidates
- Learning to See by Moving
- To Fall Or Not To Fall: A Visual Approach to Physical Stability Prediction
- A Comparative Evaluation of Approximate Probabilistic Simulation and Deep Neural Networks as Accounts of Human Physical Scene Understanding
Cited by in corpus (28)
- Explainable Machine Learning for Scientific Insights and Discoveries
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Sim4CV: A Photo-Realistic Simulator for Computer Vision Applications
- Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- Reasoning About Physical Interactions with Object-Oriented Prediction and Planning
- CoPhy: Counterfactual Learning of Physical Dynamics
- IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning
- Adaptive Neural Network-Based Approximation to Accelerate Eulerian Fluid Simulation
- Active Learning of Abstract Plan Feasibility
- Physion: Evaluating Physical Prediction from Vision in Humans and Machines
- Unsupervised Discovery of 3D Physical Objects from Video
- Bounce and Learn: Modeling Scene Dynamics with Real-World Bounces
- Forward Prediction for Physical Reasoning
- Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
- SoundSpaces: Audio-Visual Navigation in 3D Environments
- A Benchmark for Modeling Violation-of-Expectation in Physical Reasoning Across Event Categories
- Learning Long-term Visual Dynamics with Region Proposal Interaction Networks
- Learning Generalizable Physical Dynamics of 3D Rigid Objects
- Unsupervised Intuitive Physics from Past Experiences
- AVoE: A Synthetic 3D Dataset on Understanding Violation of Expectation for Artificial Cognition
- Causal World Models by Unsupervised Deconfounding of Physical Dynamics
- Structured agents for physical construction
- Adding Intuitive Physics to Neural-Symbolic Capsules Using Interaction Networks
- Predicting the Physical Dynamics of Unseen 3D Objects
- PHYRE: A New Benchmark for Physical Reasoning
- Scrutinizing and De-Biasing Intuitive Physics with Neural Stethoscopes
- Compositional Video Prediction