3 citations · 4 across the 4 of their papers we have counts for
4 papers · 1 filter
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada +4
In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce un…
Confident Approximate Policy Iteration for Efficient Local Planning in -realizable MDPs
Gellért Weisz, András György, Tadashi Kozuno +1
We consider approximate dynamic programming in -discounted Markov decision processes and apply it to approximate planning with linear value-function approximation. Our first con…
No More Pesky Hyperparameters: Offline Hyperparameter Tuning for RL
Han Wang, Archit Sakhadeo, Adam White +7
The performance of reinforcement learning (RL) agents is sensitive to the choice of hyperparameters. In real-world settings like robotics or industrial control systems, however, te…
Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences
Alan Chan, Hugo Silva, Sungsu Lim +3
Approximate Policy Iteration (API) algorithms alternate between (approximate) policy evaluation and (approximate) greedification. Many different approaches have been explored for a…