Action-Conditional Video Prediction using Deep Networks in Atari Games
arXiv:1507.08750
Abstract
Motivated by vision-based reinforcement learning (RL) problems, in particular Atari games from the recent benchmark Aracade Learning Environment (ALE), we consider spatio-temporal prediction problems where future (image-)frames are dependent on control variables or actions as well as previous frames. While not composed of natural scenes, frames in Atari games are high-dimensional in size, can involve tens of objects with one or more objects being controlled by the actions directly and many other objects being influenced indirectly, can involve entry and departure of objects, and can involve deep partial observability. We propose and evaluate two deep neural network architectures that consist of encoding, action-conditional transformation, and decoding layers based on convolutional neural networks and recurrent neural networks. Experimental results show that the proposed architectures are able to generate visually-realistic frames that are also useful for control over approximately 100-step action-conditional futures in some games. To the best of our knowledge, this paper is the first to make and evaluate long-term predictions on high-dimensional video conditioned by control inputs.
Published at NIPS 2015 (Advances in Neural Information Processing Systems 28)
References in corpus (5)
Cited by in corpus (103)
- A Brief Survey of Deep Reinforcement Learning
- Deep Reinforcement Learning for Cyber Security
- Deep Reinforcement Learning: An Overview
- VIME: Variational Information Maximizing Exploration
- Dynamic Filter Networks
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- An Algorithmic Perspective on Imitation Learning
- A Review on Deep Learning Techniques for Video Prediction
- Visual Reinforcement Learning with Imagined Goals
- Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
- Model-Ensemble Trust-Region Policy Optimization
- Unsupervised Learning for Physical Interaction through Video Prediction
- Unifying Count-Based Exploration and Intrinsic Motivation
- Learning What and Where to Draw
- Stochastic Adversarial Video Prediction
- Learning a Driving Simulator
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Emotion in Reinforcement Learning Agents and Robots: A Survey
- The Predictron: End-To-End Learning and Planning
- A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning
- A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning
- Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
- Recurrent Environment Simulators
- Self-Supervised Visual Planning with Temporal Skip Connections
- Learning to Decompose and Disentangle Representations for Video Prediction
- Learning Visual Predictive Models of Physics for Playing Billiards
- Multi-Objective Deep Reinforcement Learning
- Neural Episodic Control
- Learning Plannable Representations with Causal InfoGAN
- Dialog-based Language Learning
- Stochastic Variational Video Prediction
- Dynamic Facial Expression Generation on Hilbert Hypersphere with Conditional Wasserstein Generative Adversarial Nets
- A review of radar-based nowcasting of precipitation and applicable machine learning techniques
- Learning and Querying Fast Generative Models for Reinforcement Learning
- Learning to Perform Physics Experiments via Deep Reinforcement Learning
- Imitating Latent Policies from Observation
- Symbolic Pregression: Discovering Physical Laws from Distorted Video
- Value Prediction Network
- Generating Text with Deep Reinforcement Learning
- Independently Controllable Features
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- Visual Dynamics: Stochastic Future Generation via Layered Cross Convolutional Networks
- A Deep Learning Approach for Joint Video Frame and Reward Prediction in Atari Games
- Emergence of Communication in an Interactive World with Consistent Speakers
- The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement Learning
- Time-Agnostic Prediction: Predicting Predictable Video Frames
- Fast Generation for Convolutional Autoregressive Models
- Talking Face Generation by Conditional Recurrent Adversarial Network
- IQ of Neural Networks
- Detection and Tracking of Liquids with Fully Convolutional Networks
- Model-based Reinforcement Learning for Predictions and Control for Limit Order Books
- Deep Reinforcement Learning for Green Security Games with Real-Time Information
- Novel Video Prediction for Large-scale Scene using Optical Flow
- Mathematical Reasoning in Latent Space
- Learning and Planning with a Semantic Model
- Generalization through Simulation: Integrating Simulated and Real Data into Deep Reinforcement Learning for Vision-Based Autonomous Flight
- Vid2Game: Controllable Characters Extracted from Real-World Videos
- Flow-Grounded Spatial-Temporal Video Prediction from Still Images
- Occupancy Map Prediction Using Generative and Fully Convolutional Networks for Vehicle Navigation
- Pose Guided Human Video Generation
- Deep Learning for Reward Design to Improve Monte Carlo Tree Search in ATARI Games
- Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Representations
- Automatic Music Playlist Generation via Simulation-based Reinforcement Learning
- Model-based Behavioral Cloning with Future Image Similarity Learning
- A Neurally-Inspired Hierarchical Prediction Network for Spatiotemporal Sequence Learning and Prediction
- Model-Based Regularization for Deep Reinforcement Learning with Transcoder Networks
- Adaptive Skip Intervals: Temporal Abstraction for Recurrent Dynamical Models
- Text-to-Image Generation with Attention Based Recurrent Neural Networks
- Variational Inference for Data-Efficient Model Learning in POMDPs
- Independent Innovation Analysis for Nonlinear Vector Autoregressive Process
- Hashing over Predicted Future Frames for Informed Exploration of Deep Reinforcement Learning
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement Learning
- What Would You Do? Acting by Learning to Predict
- Over-crowdedness Alert! Forecasting the Future Crowd Distribution
- Sensorimotor Visual Perception on Embodied System Using Free Energy Principle
- Adversarial Framework for Unsupervised Learning of Motion Dynamics in Videos
- Towards a Simple Approach to Multi-step Model-based Reinforcement Learning
- ContextVP: Fully Context-Aware Video Prediction
- Moving Deep Learning into Web Browser: How Far Can We Go?
- Video Extrapolation with an Invertible Linear Embedding
- Hybrid Reinforcement Learning with Expert State Sequences
- Micro-Objective Learning : Accelerating Deep Reinforcement Learning through the Discovery of Continuous Subgoals
- Hierarchical Representation Learning for Markov Decision Processes
- Future Frame Prediction for Robot-assisted Surgery
- Deep Reinforcement Learning From Raw Pixels in Doom
- Sparse Factorization Layers for Neural Networks with Limited Supervision
- "What happens if..." Learning to Predict the Effect of Forces in Images
- NeoNav: Improving the Generalization of Visual Navigation via Generating Next Expected Observations
- Task-Relevant Object Discovery and Categorization for Playing First-person Shooter Games
- Exploiting generalization in the subspaces for faster model-based learning
- A survey of benchmarking frameworks for reinforcement learning
- Learning Representations for Pixel-based Control: What Matters and Why?
- Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
- Self-Consistent Models and Values
- High Performance Across Two Atari Paddle Games Using the Same Perceptual Control Architecture Without Training
- Planning in Dynamic Environments with Conditional Autoregressive Models
- Accelerated Target Updates for Q-learning
- Operator Shifting for Model-based Policy Evaluation
- Neural Embedding for Physical Manipulations
- Hierarchical Reinforcement Learning: Approximating Optimal Discounted TSP Using Local Policies
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- Action-conditional Sequence Modeling for Recommendation