10 citations · 10 across the 2 of their papers we have counts for
2 papers
cs.AI2021
An Analytical Update Rule for General Policy Optimization
Hepeng Li, Nicholas Clavette, Haibo He
We present an analytical policy update rule that is independent of parametric function approximators. The policy update rule is suitable for optimizing general stochastic policies…
cs.AI2020★ 10 cited
Multi-Agent Trust Region Policy Optimization
Hepeng Li, Haibo He
We extend trust region policy optimization (TRPO) to multi-agent reinforcement learning (MARL) problems. We show that the policy update of TRPO can be transformed into a distribute…