activity
20172021
most citedThe Actor-Advisor: Policy Gradient With Off-Policy Advice

3 citations · 7 across the 5 of their papers we have counts for

collaborators

7 papers

cs.AI20213 cited

Synthesising Reinforcement Learning Policies through Set-Valued Inductive Rule Learning

Youri Coppens, Denis Steckelmacher, Catholijn M. Jonker +1

Today's advanced Reinforcement Learning algorithms produce black-box policies, that are often difficult to interpret and trust for a person. We introduce a policy distilling algori…

cs.AI20191 cited

Transfer Learning Across Simulated Robots With Different Sensors

Hélène Plisnier, Denis Steckelmacher, Diederik Roijers +1

For a robot to learn a good policy, it often requires expensive equipment (such as sophisticated sensors) and a prepared training environment conducive to learning. However, it is…

cs.LG2019

Sample-Efficient Model-Free Reinforcement Learning with Off-Policy Critics

Denis Steckelmacher, Hélène Plisnier, Diederik M. Roijers +1

Value-based reinforcement-learning algorithms provide state-of-the-art results in model-free discrete-action settings, and tend to outperform actor-critic algorithms. We argue that…

cs.AI20193 cited

The Actor-Advisor: Policy Gradient With Off-Policy Advice

Hélène Plisnier, Denis Steckelmacher, Diederik M. Roijers +1

Actor-critic algorithms learn an explicit policy (actor), and an accompanying value function (critic). The actor performs actions in the environment, while the critic evaluates the…

cs.LG2018

Dynamic Weights in Multi-Objective Deep Reinforcement Learning

Axel Abels, Diederik M. Roijers, Tom Lenaerts +2

Many real-world decision problems are characterized by multiple conflicting objectives which must be balanced based on their relative importance. In the dynamic weights setting the…

cs.LG2018

Directed Policy Gradient for Safe Reinforcement Learning with Human Advice

Hélène Plisnier, Denis Steckelmacher, Tim Brys +2

Many currently deployed Reinforcement Learning agents work in an environment shared with humans, be them co-workers, users or clients. It is desirable that these agents adjust to p…