Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
arXiv:1910.10897
Abstract
Meta-reinforcement learning algorithms can enable robots to acquire new skills much more quickly, by leveraging prior experience to learn how to learn. However, much of the current research on meta-reinforcement learning focuses on task distributions that are very narrow. For example, a commonly used meta-reinforcement learning benchmark uses different running velocities for a simulated robot as different tasks. When policies are meta-trained on such narrow task distributions, they cannot possibly generalize to more quickly acquire entirely new tasks. Therefore, if the aim of these methods is to enable faster acquisition of entirely new behaviors, we must evaluate them on task distributions that are sufficiently broad to enable generalization to new behaviors. In this paper, we propose an open-source simulated benchmark for meta-reinforcement learning and multi-task learning consisting of 50 distinct robotic manipulation tasks. Our aim is to make it possible to develop algorithms that generalize to accelerate the acquisition of entirely new, held-out tasks. We evaluate 7 state-of-the-art meta-reinforcement learning and multi-task learning algorithms on these tasks. Surprisingly, while each task and its variations (e.g., with different object positions) can be learned with reasonable success, these algorithms struggle to learn with multiple tasks at the same time, even with as few as ten distinct training tasks. Our analysis and open-source environments pave the way for future research in multi-task learning and meta-learning that can enable meaningful generalization, thereby unlocking the full potential of these methods.
This is an update version of a manuscript that originally appeared at CoRL 2019. Videos are here: meta-world.github.io, open-sourced code are available at: https://github.com/rlworkgroup/metaworld, and the baselines can be found at https://github.com/rlworkgroup/garage
References in corpus (28)
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Trust Region Policy Optimization
- A Simple Neural Attentive Meta-Learner
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- DeepMind Control Suite
- Learning to reinforcement learn
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Visual Reinforcement Learning with Imagined Goals
- CARLA: An Open Urban Driving Simulator
- Learning Dexterous In-Hand Manipulation
- Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
- Quantifying Generalization in Reinforcement Learning
- MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments
- Meta Reinforcement Learning with Latent Variable Gaussian Processes
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
- Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning
- Gotta Learn Fast: A New Benchmark for Generalization in RL
- HoME: a Household Multimodal Environment
- Robot Learning in Homes: Improving Generalization and Reducing Dataset Bias
- RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation
- Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
- Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation
- Been There, Done That: Meta-Learning with Episodic Recall
- Meta-Learning by the Baldwin Effect
- Multi-task Deep Reinforcement Learning with PopArt
- ProMP: Proximal Meta-Policy Search
Cited by in corpus (107)
- Multi-Task Learning with Deep Neural Networks: A Survey
- Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments
- dm_control: Software and Tasks for Continuous Control
- A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
- Leveraging Procedural Generation to Benchmark Reinforcement Learning
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
- Gradient Surgery for Multi-Task Learning
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- MERLIN: Multi-agent offline and transfer learning for occupant-centric energy flexible operation of grid-interactive communities using smart meter data and CityLearn
- Multi-Task Reinforcement Learning with Soft Modularization
- Avoiding Catastrophe: Active Dendrites Enable Multi-Task Learning in Dynamic Environments
- COMBO: Conservative Offline Model-Based Policy Optimization
- Meta-Learning Requires Meta-Augmentation
- AllenAct: A Framework for Embodied AI Research
- iGibson 1.0: a Simulation Environment for Interactive Tasks in Large Realistic Scenes
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
- From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence
- CausalWorld: A Robotic Manipulation Benchmark for Causal Structure and Transfer Learning
- Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
- Maximum Entropy RL (Provably) Solves Some Robust RL Problems
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Multi-Task Reinforcement Learning with Context-based Representations
- Generalized Hindsight for Reinforcement Learning
- Meta-learning curiosity algorithms
- Offline Reinforcement Learning from Images with Latent Space Models
- OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning
- Learning to design without prior data: Discovering generalizable design strategies using deep learning and tree search
- Learning Visual Robotic Control Efficiently with Contrastive Pre-training and Data Augmentation
- Analysis of Randomization Effects on Sim2Real Transfer in Reinforcement Learning for Robotic Manipulation Tasks
- Online Fast Adaptation and Knowledge Accumulation: a New Approach to Continual Learning
- Learning Object Manipulation Skills via Approximate State Estimation from Real Videos
- PixL2R: Guiding Reinforcement Learning Using Natural Language by Mapping Pixels to Rewards
- Robust Predictable Control
- Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
- Procedural Content Generation: Better Benchmarks for Transfer Reinforcement Learning
- Challenges of Applying Deep Reinforcement Learning in Dynamic Dispatching
- Few-shot Quality-Diversity Optimization
- Offline Meta-Reinforcement Learning with Advantage Weighting
- Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without Forgetting
- Evaluating Agents without Rewards
- WordCraft: An Environment for Benchmarking Commonsense Agents
- Context Meta-Reinforcement Learning via Neuromodulation
- CORA: Benchmarks, Baselines, and Metrics as a Platform for Continual Reinforcement Learning Agents
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration
- Modeling and Optimization Trade-off in Meta-learning
- Curriculum in Gradient-Based Meta-Reinforcement Learning
- Zero-Shot Terrain Generalization for Visual Locomotion Policies
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices
- The MAGICAL Benchmark for Robust Imitation
- Measuring and Harnessing Transference in Multi-Task Learning
- rl_reach: Reproducible Reinforcement Learning Experiments for Robotic Reaching Tasks
- Powerpropagation: A sparsity inducing weight reparameterisation
- Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis
- When Is Generalizable Reinforcement Learning Tractable?
- A Study of Continual Learning Methods for Q-Learning
- Systematic Evaluation of Causal Discovery in Visual Model Based Reinforcement Learning
- S4RL: Surprisingly Simple Self-Supervision for Offline Reinforcement Learning
- Reinforcement Learning with Videos: Combining Offline Observations with Interaction
- Robust Domain Randomised Reinforcement Learning through Peer-to-Peer Distillation
- Advances in Preference-based Reinforcement Learning: A Review
- Learning Multi-Task Transferable Rewards via Variational Inverse Reinforcement Learning
- A Risk-Sensitive Approach to Policy Optimization
- Meta-Model-Based Meta-Policy Optimization
- Meta-Reinforcement Learning in Broad and Non-Parametric Environments
- PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation
- Replacing Rewards with Examples: Example-Based Policy Search via Recursive Classification
- Procedural Generalization by Planning with Self-Supervised World Models
- Unsupervised Visual Attention and Invariance for Reinforcement Learning
- Sequoia: A Software Framework to Unify Continual Learning Research
- A Channel Coding Benchmark for Meta-Learning
- CARL: A Benchmark for Contextual and Adaptive Reinforcement Learning
- Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning
- Scenic4RL: Programmatic Modeling and Generation of Reinforcement Learning Environments
- Adaptive Procedural Task Generation for Hard-Exploration Problems
- T3S: Improving Multi-Task Reinforcement Learning with Task-Specific Feature Selector and Scheduler
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
- Bayes meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
- CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP
- Touch-based Curiosity for Sparse-Reward Tasks
- Improving Context-Based Meta-Reinforcement Learning with Self-Supervised Trajectory Contrastive Learning
- Hindsight Foresight Relabeling for Meta-Reinforcement Learning
- Bias-reduced Multi-step Hindsight Experience Replay for Efficient Multi-goal Reinforcement Learning
- Off-Policy Meta-Reinforcement Learning Based on Feature Embedding Spaces
- A Brief Look at Generalization in Visual Meta-Reinforcement Learning
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement Learning
- Meta Arcade: A Configurable Environment Suite for Meta-Learning
- Context-Based Soft Actor Critic for Environments with Non-stationary Dynamics
- What is Going on Inside Recurrent Meta Reinforcement Learning Agents?
- Koopman Q-learning: Offline Reinforcement Learning via Symmetries of Dynamics
- Connecting Context-specific Adaptation in Humans to Meta-learning
- AMaze: An intuitive benchmark generator for fast prototyping of generalizable agents
- Discovering Generalizable Skills via Automated Generation of Diverse Tasks
- Reinforcement Learning in the Wild with Maximum Likelihood-based Model Transfer
- Complex Skill Acquisition Through Simple Skill Imitation Learning
- Learning Transferable Concepts in Deep Reinforcement Learning
- Management of Resource at the Network Edge for Federated Learning
- Assisted Robust Reward Design
- gym-saturation: Gymnasium environments for saturation provers (System description)
- Suboptimal coverings for continuous spaces of control tasks
- Visual Goal-Directed Meta-Learning with Contextual Planning Networks
- CAZSL: Zero-Shot Regression for Pushing Models by Generalizing Through Context
- Adaptation of Quadruped Robot Locomotion with Meta-Learning
- When Autonomous Systems Meet Accuracy and Transferability through AI: A Survey
- Learning a Skill-sequence-dependent Policy for Long-horizon Manipulation Tasks
- Lifetime policy reuse and the importance of task capacity