Imitating and Finetuning Model Predictive Control for Robust and Symmetric Quadrupedal Locomotion
arXiv:2311.02304 · doi:10.1109/LRA.2023.3320827
Abstract
Control of legged robots is a challenging problem that has been investigated by different approaches, such as model-based control and learning algorithms. This work proposes a novel Imitating and Finetuning Model Predictive Control (IFM) framework to take the strengths of both approaches. Our framework first develops a conventional model predictive controller (MPC) using Differential Dynamic Programming and Raibert heuristic, which serves as an expert policy. Then we train a clone of the MPC using imitation learning to make the controller learnable. Finally, we leverage deep reinforcement learning with limited exploration for further finetuning the policy on more challenging terrains. By conducting comprehensive simulation and hardware experiments, we demonstrate that the proposed IFM framework can significantly improve the performance of the given MPC controller on rough, slippery, and conveyor terrains that require careful coordination of footsteps. We also showcase that IFM can efficiently produce more symmetric, periodic, and energy-efficient gaits compared to Vanilla RL with a minimal burden of reward shaping.
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Learning agile and dynamic motor skills for legged robots
- Learning Quadrupedal Locomotion over Challenging Terrain
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
- Learning robust perceptive locomotion for quadrupedal robots in the wild
- Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control
- Multi-expert learning of adaptive legged locomotion
- Concurrent Training of a Control Policy and a State Estimator for Dynamic and Robust Legged Locomotion
- Representation-Free Model Predictive Control for Dynamic Motions in Quadrupeds
- Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning
- MPC-Net: A First Principles Guided Policy Search
- Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
Cited by in corpus (3)
- Contact-Implicit Model Predictive Control: Controlling Diverse Quadruped Motions Without Pre-Planned Contact Modes or Trajectories
- Designing a skilled soccer team for RoboCup: exploring skill-set-primitives through reinforcement learning
- PPF: Pre-training and Preservative Fine-tuning of Humanoid Locomotion via Model-Assumption-based Regularization