activity
20182022
most citedAn Autonomous Free Airspace En-route Controller using Deep Reinforcement Learning Techniques

18 citations · 32 across the 9 of their papers we have counts for

collaborators

14 papers

cs.AI20223 cited

Logic-based AI for Interpretable Board Game Winner Prediction with Tsetlin Machine

Charul Giri, Ole-Christoffer Granmo, Herke van Hoof +1

Hex is a turn-based two-player connection game with a high branching factor, making the game arbitrarily complex with increasing board sizes. As such, top-performing algorithms for…

cs.AI2022

Reliably Re-Acting to Partner's Actions with the Social Intrinsic Motivation of Transfer Empowerment

Tessa van der Heiden, Herke van Hoof, Efstratios Gavves +1

We consider multi-agent reinforcement learning (MARL) for cooperative communication and coordination tasks. MARL agents can be brittle because they can overfit their training partn…

cs.LG2022

Fast and Data Efficient Reinforcement Learning from Pixels via Non-Parametric Value Approximation

Alexander Long, Alan Blair, Herke van Hoof

We present Nonparametric Approximation of Inter-Trace returns (NAIT), a Reinforcement Learning algorithm for discrete action, pixel-based environments that is both highly sample an…

cs.LG20215 cited

A Survey of Exploration Methods in Reinforcement Learning

Susan Amin, Maziar Gomrokchi, Harsh Satija +2

Exploration is an essential component of reinforcement learning algorithms, where agents need to learn how to predict and control unknown and often stochastic environments. Reinfor…

cs.LG2021

Combining Reward Information from Multiple Sources

Dmitrii Krasheninnikov, Rohin Shah, Herke van Hoof

Given two sources of evidence about a latent variable, one can combine the information from both by multiplying the likelihoods of each piece of evidence. However, when one or both…

cs.LG20211 cited

Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models

Qi Wang, Herke van Hoof

Reinforcement learning is a promising paradigm for solving sequential decision-making problems, but low data efficiency and weak generalization across tasks are bottlenecks in real…