End-to-End Training of Deep Visuomotor Policies
arXiv:1504.00702
Abstract
Policy search methods can allow robots to learn control policies for a wide range of tasks, but practical applications of policy search often require hand-engineered components for perception, state estimation, and low-level control. In this paper, we aim to answer the following question: does training the perception and control systems jointly end-to-end provide better performance than training each component separately? To this end, we develop a method that can be used to learn policies that map raw image observations directly to torques at the robot's motors. The policies are represented by deep convolutional neural networks (CNNs) with 92,000 parameters, and are trained using a partially observed guided policy search method, which transforms policy search into supervised learning, with supervision provided by a simple trajectory-centric reinforcement learning method. We evaluate our method on a range of real-world manipulation tasks that require close coordination between vision and control, such as screwing a cap onto a bottle, and present simulated comparisons to a range of prior policy search methods.
updating with revisions for JMLR final version
References in corpus (8)
- Deep Learning in Neural Networks: An Overview
- Continuous control with deep reinforcement learning
- Caffe: Convolutional Architecture for Fast Feature Embedding
- On the difficulty of training Recurrent Neural Networks
- Going Deeper with Convolutions
- Path Integral Policy Improvement with Covariance Matrix Adaptation
- Learning Contact-Rich Manipulation Skills with Guided Policy Search
- Supersizing Self-supervision: Learning to Grasp from 50K Tries and 700 Robot Hours
Cited by in corpus (162)
- Continuous control with deep reinforcement learning
- A Brief Survey of Deep Reinforcement Learning
- Dueling Network Architectures for Deep Reinforcement Learning
- Asynchronous Methods for Deep Reinforcement Learning
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Emergence of Locomotion Behaviours in Rich Environments
- SeqSleepNet: End-to-End Hierarchical Recurrent Neural Network for Sequence-to-Sequence Automatic Sleep Staging
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Deep Exploration via Bootstrapped DQN
- Reinforcement Learning with Deep Energy-Based Policies
- An Algorithmic Perspective on Imitation Learning
- Hindsight Experience Replay
- GCN-RL Circuit Designer: Transferable Transistor Sizing with Graph Neural Networks and Reinforcement Learning
- Learning to Predict the Cosmological Structure Formation
- Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
- Memory-based control with recurrent neural networks
- Sample Efficient Actor-Critic with Experience Replay
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
- Flow: A Modular Learning Framework for Mixed Autonomy Traffic
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Data-Driven Control of Complex Networks
- Machine Learning for Intelligent Optical Networks: A Comprehensive Survey
- Learning Force Control for Contact-rich Manipulation Tasks with Rigid Position-controlled Robots
- Learning to Optimize
- Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense
- Collective Robot Reinforcement Learning with Distributed Asynchronous Guided Policy Search
- Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks
- Data-efficient Deep Reinforcement Learning for Dexterous Manipulation
- Reinforcement Learning Neural Turing Machines - Revised
- Reinforcement and Imitation Learning for Diverse Visuomotor Skills
- Learning Visual Predictive Models of Physics for Playing Billiards
- Learning and Transfer of Modulated Locomotor Controllers
- Multi-Objective Deep Reinforcement Learning
- Combining policy gradient and Q-learning
- Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning
- Learning Plannable Representations with Causal InfoGAN
- Universal Planning Networks
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks
- Strategic Attentive Writer for Learning Macro-Actions
- Machine learning for industrial sensing and control: A survey and practical perspective
- Learning to Optimize Neural Nets
- Vision-based Driver Assistance Systems: Survey, Taxonomy and Advances
- Intrinsically motivated reinforcement learning for human-robot interaction in the real-world
- Vision-driven Compliant Manipulation for Reliable, High-Precision Assembly Tasks
- End-to-End Pixel-Based Deep Active Inference for Body Perception and Action
- Deep reinforcement learning for guidewire navigation in coronary artery phantom
- Neural network-based clustering using pairwise constraints
- DPC-Net: Deep Pose Correction for Visual Localization
- Accelerating Reinforcement Learning for Reaching using Continuous Curriculum Learning
- Deep Reinforcement Learning in Parameterized Action Space
- A Survey on Physics Informed Reinforcement Learning: Review and Open Problems
- Challenging Machine Learning-based Clone Detectors via Semantic-preserving Code Transformations
- Towards human-level performance on automatic pose estimation of infant spontaneous movements
- Transformer-based Model Predictive Control: Trajectory Optimization via Sequence Modeling
- Are we done with object recognition? The iCub robot's perspective
- Modern Machine Learning Tools for Monitoring and Control of Industrial Processes: A Survey
- Heatmap Regression via Randomized Rounding
- Diffusion-based neuromodulation can eliminate catastrophic forgetting in simple neural networks
- Semi-Automatic Data Annotation guided by Feature Space Projection
- Learning Dexterous Manipulation Policies from Experience and Imitation
- Deep Q-Network Based Decision Making for Autonomous Driving
- DRLViz: Understanding Decisions and Memory in Deep Reinforcement Learning
- Visual Spatial Attention and Proprioceptive Data-Driven Reinforcement Learning for Robust Peg-in-Hole Task Under Variable Conditions
- Memory Augmented Control Networks
- Learning Dexterous Manipulation for a Soft Robotic Hand from Human Demonstration
- Gaze-based dual resolution deep imitation learning for high-precision dexterous robot manipulation
- Ultrasound Image Representation Learning by Modeling Sonographer Visual Attention
- Uncovering Instabilities in Variational-Quantum Deep Q-Networks
- Supersizing Self-supervision: Learning to Grasp from 50K Tries and 700 Robot Hours
- Learning Robotic Navigation from Experience: Principles, Methods, and Recent Results
- Least-Squares Temporal Difference Learning for the Linear Quadratic Regulator
- HIRL: Hierarchical Inverse Reinforcement Learning for Long-Horizon Tasks with Delayed Rewards
- Learning Transferable Policies for Monocular Reactive MAV Control
- How to Train a CAT: Learning Canonical Appearance Transformations for Direct Visual Localization Under Illumination Change
- Motion Switching with Sensory and Instruction Signals by designing Dynamical Systems using Deep Neural Network
- Trust-Region Method with Deep Reinforcement Learning in Analog Design Space Exploration
- Deep reinforcement learning in World-Earth system models to discover sustainable management strategies
- Learning Event-triggered Control from Data through Joint Optimization
- Neural radiance fields in the industrial and robotics domain: applications, research opportunities and use cases
- Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics Model
- Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation
- Data-Efficient Learning of Feedback Policies from Image Pixels using Deep Dynamical Models
- Graph Policy Gradients for Large Scale Robot Control
- BatMobility: Towards Flying Without Seeing for Autonomous Drones
- A Data-Efficient Deep Learning Approach for Deployable Multimodal Social Robots
- Compatible Value Gradients for Reinforcement Learning of Continuous Deep Policies
- Value Iteration Networks on Multiple Levels of Abstraction
- Distributed multi-agent target search and tracking with Gaussian process and reinforcement learning
- Deep Reinforcement Learning for Dynamic Multichannel Access in Wireless Networks
- Learning Deep Control Policies for Autonomous Aerial Vehicles with MPC-Guided Policy Search
- Learning a Low-dimensional Representation of a Safe Region for Safe Reinforcement Learning on Dynamical Systems
- Quantile QT-Opt for Risk-Aware Vision-Based Robotic Grasping
- Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation
- Learning Connectivity-Maximizing Network Configurations
- Behavior Self-Organization Supports Task Inference for Continual Robot Learning
- Quantum Bandits
- Map-based Experience Replay: A Memory-Efficient Solution to Catastrophic Forgetting in Reinforcement Learning
- Continual Robot Learning using Self-Supervised Task Inference
- Integrating Contrastive Learning with Dynamic Models for Reinforcement Learning from Images
- Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination
- Patterns for Learning with Side Information
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation
- Multi-Objective Convolutional Neural Networks for Robot Localisation and 3D Position Estimation in 2D Camera Images
- Modelling Generalized Forces with Reinforcement Learning for Sim-to-Real Transfer
- Deep Model Predictive Control with Stability Guarantees
- Hindsight Goal Ranking on Replay Buffer for Sparse Reward Environment
- Scalable Centralized Deep Multi-Agent Reinforcement Learning via Policy Gradients
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Detection and Tracking of Liquids with Fully Convolutional Networks
- Learning Visual Robotic Control Efficiently with Contrastive Pre-training and Data Augmentation
- GoSafeOpt: Scalable Safe Exploration for Global Optimization of Dynamical Systems
- Learning to Navigate Using Mid-Level Visual Priors
- Cloud-Edge Training Architecture for Sim-to-Real Deep Reinforcement Learning
- Combining Model-Based and Model-Free Methods for Nonlinear Control: A Provably Convergent Policy Gradient Approach
- From explanation to synthesis: Compositional program induction for learning from demonstration
- One-Shot Object Localization Using Learnt Visual Cues via Siamese Networks
- A Survey of Behavior Learning Applications in Robotics -- State of the Art and Perspectives
- Automated Design and Optimization of Distributed Filtering Circuits via Reinforcement Learning
- DyNODE: Neural Ordinary Differential Equations for Dynamics Modeling in Continuous Control
- Learning Deep Neural Network Policies with Continuous Memory States
- The PlayStation Reinforcement Learning Environment (PSXLE)
- Learning to Control DC Motor for Micromobility in Real Time with Reinforcement Learning
- Autonomous Marker-less Rapid Aerial Grasping
- Learning Relevant Features for Manipulation Skills using Meta-Level Priors
- Reinforced Labels: Multi-Agent Deep Reinforcement Learning for Point-Feature Label Placement
- Pittsburgh Learning Classifier Systems for Explainable Reinforcement Learning: Comparing with XCS
- Towards Practical Implementations of Person Re-Identification from Full Video Frames
- Policy Gradients for Probabilistic Constrained Reinforcement Learning
- Human-like Clustering with Deep Convolutional Neural Networks
- FOSS: A Self-Learned Doctor for Query Optimizer
- Seeing All the Angles: Learning Multiview Manipulation Policies for Contact-Rich Tasks from Demonstrations
- Multi Agent Navigation in Unconstrained Environments using a Centralized Attention based Graphical Neural Network Controller
- Deep Robot Sketching: An application of Deep Q-Learning Networks for human-like sketching
- Safe Learning MPC with Limited Model Knowledge and Data
- Teaching Robots Novel Objects by Pointing at Them
- On the Existence and Computation of Minimum Attention Optimal Control Laws
- Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning
- Enhancing a Neurocognitive Shared Visuomotor Model for Object Identification, Localization, and Grasping With Learning From Auxiliary Tasks
- Neural Style Transfer with Twin-Delayed DDPG for Shared Control of Robotic Manipulators
- Deep3D: Fully Automatic 2D-to-3D Video Conversion with Deep Convolutional Neural Networks
- Visual-Policy Learning through Multi-Camera View to Single-Camera View Knowledge Distillation for Robot Manipulation Tasks
- Dynamically Feasible Deep Reinforcement Learning Policy for Robot Navigation in Dense Mobile Crowds
- Synthesis-guided Adversarial Scenario Generation for Gray-box Feedback Control Systems with Sensing Imperfections
- Guided Policy Search using Sequential Convex Programming for Initialization of Trajectory Optimization Algorithms
- MobILE: Model-Based Imitation Learning From Observation Alone
- Gaussian Vector: An Efficient Solution for Facial Landmark Detection
- Expert Level control of Ramp Metering based on Multi-task Deep Reinforcement Learning
- Seamless Integration and Coordination of Cognitive Skills in Humanoid Robots: A Deep Learning Approach
- Discovering Latent States for Model Learning: Applying Sensorimotor Contingencies Theory and Predictive Processing to Model Context
- Investigation of Factorized Optical Flows as Mid-Level Representations
- Policy Optimization with Second-Order Advantage Information
- Skeleton Transformer Networks: 3D Human Pose and Skinned Mesh from Single RGB Image
- Learning Fast and Precise Pixel-to-Torque Control
- JAX-IK: Real-Time Inverse Kinematics for Generating Multi-Constrained Movements of Virtual Human Characters
- Learning in Sparse Rewards settings through Quality-Diversity algorithms
- Room Clearance with Feudal Hierarchical Reinforcement Learning
- Interpretable Option Discovery using Deep Q-Learning and Variational Autoencoders
- Active Object Localization with Deep Reinforcement Learning
- Associating Grasp Configurations with Hierarchical Features in Convolutional Neural Networks
- Inferring the Optimal Policy using Markov Chain Monte Carlo