747 citations · 782 across the 14 of their papers we have counts for
3 papers · 1 filter
Quantile Filtered Imitation Learning
David Brandfonbrener, William F. Whitney, Rajesh Ranganath +1
We introduce quantile filtered imitation learning (QFIL), a novel policy improvement operator designed for offline reinforcement learning. QFIL performs policy improvement by runni…
Offline RL Without Off-Policy Evaluation
David Brandfonbrener, William F. Whitney, Rajesh Ranganath +1
Most prior approaches to offline reinforcement learning (RL) have taken an iterative actor-critic approach involving off-policy evaluation. In this paper we show that simply doing…
Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning
William F. Whitney, Michael Bloesch, Jost Tobias Springenberg +3
Despite the close connection between exploration and sample efficiency, most state of the art reinforcement learning algorithms include no considerations for exploration beyond max…