activity
20162020
most citedSafe Exploration in Continuous Action Spaces

275 citations · 286 across the 4 of their papers we have counts for

collaborators

8 papers

cs.LG20203 cited

The Architectural Implications of Distributed Reinforcement Learning on CPU-GPU Systems

Ahmet Inci, Evgeny Bolotin, Yaosheng Fu +4

With deep reinforcement learning (RL) methods achieving results that exceed human capabilities in games, robotics, and simulated environments, continued scaling of RL training is c…

cs.LG20196 cited

A Tale of Two-Timescale Reinforcement Learning with the Tightest Finite-Time Bound

Gal Dalal, Balazs Szorenyi, Gugan Thoppe

Policy evaluation in reinforcement learning is often conducted using two-timescale stochastic approximation, which results in various gradient temporal difference methods such as G…

cs.LG2018

How to Combine Tree-Search Methods in Reinforcement Learning

Yonathan Efroni, Gal Dalal, Bruno Scherrer +1

Finite-horizon lookahead policies are abundantly used in Reinforcement Learning and demonstrate impressive empirical success. Usually, the lookahead policies are implemented with s…

cs.LG2018

Multiple-Step Greedy Policies in Online and Approximate Reinforcement Learning

Yonathan Efroni, Gal Dalal, Bruno Scherrer +1

Multiple-step lookahead policies have demonstrated high empirical competence in Reinforcement Learning, via the use of Monte Carlo Tree Search or Model Predictive Control. In a rec…

cs.AI2018

Beyond the One Step Greedy Approach in Reinforcement Learning

Yonathan Efroni, Gal Dalal, Bruno Scherrer +1

The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation…

cs.AI2018275 cited

Safe Exploration in Continuous Action Spaces

Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik +3

We address the problem of deploying a reinforcement learning (RL) agent on a physical system such as a datacenter cooling unit or robot, where critical constraints must never be vi…