Robotic Table Tennis: A Case Study into a High Speed Learning System
arXiv:2309.03315 · doi:10.15607/RSS.2023.XIX.006
Abstract
We present a deep-dive into a real-world robotic learning system that, in previous work, was shown to be capable of hundreds of table tennis rallies with a human and has the ability to precisely return the ball to desired targets. This system puts together a highly optimized perception subsystem, a high-speed low-latency robot controller, a simulation paradigm that can prevent damage in the real world and also train policies for zero-shot transfer, and automated real world environment resets that enable autonomous training and evaluation on physical robots. We complement a complete system description, including numerous design decisions that are typically not widely disseminated, with a collection of studies that clarify the importance of mitigating various sources of latency, accounting for training and deployment distribution shifts, robustness of the perception system, sensitivity to policy hyper-parameters, and choice of action space. A video demonstrating the components of the system and details of experimental results can be found at https://youtu.be/uFcnWjB42I0.
Published and presented at Robotics: Science and Systems (RSS2023)
References in corpus (24)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- End to End Learning for Self-Driving Cars
- Learning agile and dynamic motor skills for legged robots
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Sim-to-Real Transfer of Robotic Control with Dynamics Randomization
- Solving Rubik's Cube with a Robot Hand
- OpenAI Gym
- Glow: Generative Flow with Invertible 1x1 Convolutions
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- Structured Evolution with Compact Architectures for Scalable Policy Optimization
- Learning from Suboptimal Demonstration via Self-Supervised Reward Regression
- Legged Locomotion in Challenging Terrains using Egocentric Vision
- A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning
- Optimal Stroke Learning with Policy Gradient Approach for Robotic Table Tennis
- i-Sim2Real: Reinforcement Learning of Robotic Policies in Tight Human-Robot Interaction Loops
- Joint Goal and Strategy Inference across Heterogeneous Demonstrators via Reward Network Distillation
- Learning to Play Table Tennis From Scratch using Muscular Robots
- Local Metrics for Multi-Object Tracking
- Robotic Table Tennis with Model-Free Reinforcement Learning
- Safe Reinforcement Learning for Legged Locomotion
- GoalsEye: Learning High Speed Precision Table Tennis on a Physical Robot
- Agile Catching with Whole-Body MPC and Blackbox Policy Learning