275 citations · 286 across the 4 of their papers we have counts for
8 papers
The Architectural Implications of Distributed Reinforcement Learning on CPU-GPU Systems
Ahmet Inci, Evgeny Bolotin, Yaosheng Fu +4
With deep reinforcement learning (RL) methods achieving results that exceed human capabilities in games, robotics, and simulated environments, continued scaling of RL training is c…
A Tale of Two-Timescale Reinforcement Learning with the Tightest Finite-Time Bound
Gal Dalal, Balazs Szorenyi, Gugan Thoppe
Policy evaluation in reinforcement learning is often conducted using two-timescale stochastic approximation, which results in various gradient temporal difference methods such as G…
How to Combine Tree-Search Methods in Reinforcement Learning
Yonathan Efroni, Gal Dalal, Bruno Scherrer +1
Finite-horizon lookahead policies are abundantly used in Reinforcement Learning and demonstrate impressive empirical success. Usually, the lookahead policies are implemented with s…
Multiple-Step Greedy Policies in Online and Approximate Reinforcement Learning
Yonathan Efroni, Gal Dalal, Bruno Scherrer +1
Multiple-step lookahead policies have demonstrated high empirical competence in Reinforcement Learning, via the use of Monte Carlo Tree Search or Model Predictive Control. In a rec…
Beyond the One Step Greedy Approach in Reinforcement Learning
Yonathan Efroni, Gal Dalal, Bruno Scherrer +1
The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation…
Safe Exploration in Continuous Action Spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik +3
We address the problem of deploying a reinforcement learning (RL) agent on a physical system such as a datacenter cooling unit or robot, where critical constraints must never be vi…