14 citations · 43 across the 12 of their papers we have counts for
5 papers · 1 filter
Monte Carlo Tree Search Algorithms for Risk-Aware and Multi-Objective Reinforcement Learning
Conor F. Hayes, Mathieu Reymond, Diederik M. Roijers +2
In many risk-aware and multi-objective reinforcement learning settings, the utility of the user is derived from a single execution of a policy. In these settings, making decisions…
Transfer Learning Across Simulated Robots With Different Sensors
Hélène Plisnier, Denis Steckelmacher, Diederik Roijers +1
For a robot to learn a good policy, it often requires expensive equipment (such as sophisticated sensors) and a prepared training environment conducive to learning. However, it is…
The Actor-Advisor: Policy Gradient With Off-Policy Advice
Hélène Plisnier, Denis Steckelmacher, Diederik M. Roijers +1
Actor-critic algorithms learn an explicit policy (actor), and an accompanying value function (critic). The actor performs actions in the environment, while the critic evaluates the…
Reinforcement Learning in POMDPs with Memoryless Options and Option-Observation Initiation Sets
Denis Steckelmacher, Diederik M. Roijers, Anna Harutyunyan +3
Many real-world reinforcement learning problems have a hierarchical nature, and often exhibit some degree of partial observability. While hierarchy and partial observability are us…
Structure in the Value Function of Two-Player Zero-Sum Games of Incomplete Information
Auke J. Wiggers, Frans A. Oliehoek, Diederik M. Roijers
Zero-sum stochastic games provide a rich model for competitive decision making. However, under general forms of state uncertainty as considered in the Partially Observable Stochast…