activity
20142019
most citedOff-Policy Shaping Ensembles in Reinforcement Learning

4 citations · 8 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI20191 cited

Transfer Learning Across Simulated Robots With Different Sensors

Hélène Plisnier, Denis Steckelmacher, Diederik Roijers +1

For a robot to learn a good policy, it often requires expensive equipment (such as sophisticated sensors) and a prepared training environment conducive to learning. However, it is…

cs.LG2019

Sample-Efficient Model-Free Reinforcement Learning with Off-Policy Critics

Denis Steckelmacher, Hélène Plisnier, Diederik M. Roijers +1

Value-based reinforcement-learning algorithms provide state-of-the-art results in model-free discrete-action settings, and tend to outperform actor-critic algorithms. We argue that…

cs.AI20193 cited

The Actor-Advisor: Policy Gradient With Off-Policy Advice

Hélène Plisnier, Denis Steckelmacher, Diederik M. Roijers +1

Actor-critic algorithms learn an explicit policy (actor), and an accompanying value function (critic). The actor performs actions in the environment, while the critic evaluates the…

cs.AI2017

Reinforcement Learning in POMDPs with Memoryless Options and Option-Observation Initiation Sets

Denis Steckelmacher, Diederik M. Roijers, Anna Harutyunyan +3

Many real-world reinforcement learning problems have a hierarchical nature, and often exhibit some degree of partial observability. While hierarchy and partial observability are us…

cs.AI20144 cited

Off-Policy Shaping Ensembles in Reinforcement Learning

Anna Harutyunyan, Tim Brys, Peter Vrancx +1

Recent advances of gradient temporal-difference methods allow to learn off-policy multiple value functions in parallel with- out sacrificing convergence guarantees or computational…