PFPN: Continuous Control of Physically Simulated Characters using Particle Filtering Policy Network
arXiv:2003.06959 · doi:10.1145/3487983.3488301
Abstract
Data-driven methods for physics-based character control using reinforcement learning have been successfully applied to generate high-quality motions. However, existing approaches typically rely on Gaussian distributions to represent the action policy, which can prematurely commit to suboptimal actions when solving high-dimensional continuous control problems for highly-articulated characters. In this paper, to improve the learning performance of physics-based character controllers, we propose a framework that considers a particle-based action policy as a substitute for Gaussian policies. We exploit particle filtering to dynamically explore and discretize the action space, and track the posterior policy represented as a mixture distribution. The resulting policy can replace the unimodal Gaussian policy which has been the staple for character control problems, without changing the underlying model architecture of the reinforcement learning algorithm used to perform policy optimization. We demonstrate the applicability of our approach on various motion capture imitation tasks. Baselines using our particle-based policies achieve better imitation performance and speed of convergence as compared to corresponding implementations using Gaussians, and are more robust to external perturbations during character control. Related code is available at: https://motion-lab.github.io/PFPN.
Motion, Interaction and Games (MIG '21)
References in corpus (13)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Addressing Function Approximation Error in Actor-Critic Methods
- Soft Actor-Critic Algorithms and Applications
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
- Emergence of Locomotion Behaviours in Rich Environments
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- MoGlow: Probabilistic and controllable motion synthesis using normalising flows
- Synthesis of Biologically Realistic Human Motion Using Joint Torque Actuation
- MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies
- Learning to Locomote: Understanding How Environment Design Matters for Deep Reinforcement Learning
- Boosting Trust Region Policy Optimization by Normalizing Flows Policy
- Discretizing Continuous Action Space for On-Policy Optimization