18 citations · 43 across the 15 of their papers we have counts for
5 papers · 1 filter
Calculus on MDPs: Potential Shaping as a Gradient
Erik Jenner, Herke van Hoof, Adam Gleave
In reinforcement learning, different reward functions can be equivalent in terms of the optimal policies they induce. A particularly well-known and important example is potential s…
Neural Topological Ordering for Computation Graphs
Mukul Gagrani, Corrado Rainone, Yang Yang +7
Recent works on machine learning for combinatorial optimization have shown that learning based approaches can outperform heuristic methods in terms of speed and performance. In thi…
Logic-based AI for Interpretable Board Game Winner Prediction with Tsetlin Machine
Charul Giri, Ole-Christoffer Granmo, Herke van Hoof +1
Hex is a turn-based two-player connection game with a high branching factor, making the game arbitrarily complex with increasing board sizes. As such, top-performing algorithms for…
Reliably Re-Acting to Partner's Actions with the Social Intrinsic Motivation of Transfer Empowerment
Tessa van der Heiden, Herke van Hoof, Efstratios Gavves +1
We consider multi-agent reinforcement learning (MARL) for cooperative communication and coordination tasks. MARL agents can be brittle because they can overfit their training partn…
Fast and Data Efficient Reinforcement Learning from Pixels via Non-Parametric Value Approximation
Alexander Long, Alan Blair, Herke van Hoof
We present Nonparametric Approximation of Inter-Trace returns (NAIT), a Reinforcement Learning algorithm for discrete action, pixel-based environments that is both highly sample an…