2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Erinc Merdivan, Sten Hanke, Matthieu Geist
Recent successful deep reinforcement learning algorithms, such as Trust Region Policy Optimization (TRPO) or Proximal Policy Optimization (PPO), are fundamentally variations of con…