7 citations · 7 across the 1 of their papers we have counts for
1 paper · 1 filter
David Pfau, Ian Davies, Diana Borsa +3
We introduce Wasserstein Policy Optimization (WPO), an actor-critic algorithm for reinforcement learning in continuous action spaces. WPO can be derived as an approximation to Wass…