187 citations · 193 across the 3 of their papers we have counts for
9 papers
Deep Reinforcement Learning for Navigation in AAA Video Games
Eloi Alonso, Maxim Peter, David Goumard +1
In video games, non-player characters (NPCs) are used to enhance the players' experience in a variety of ways, e.g., as enemies, allies, or innocent bystanders. A crucial component…
TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning?
Joshua Romoff, Peter Henderson, David Kanaa +4
We investigate whether Jacobi preconditioning, accounting for the bootstrap term in temporal difference (TD) learning, can help boost performance of adaptive optimizers. Our method…
Gossip-based Actor-Learner Architectures for Deep Reinforcement Learning
Mahmoud Assran, Joshua Romoff, Nicolas Ballas +2
Multi-simulator training has contributed to the recent success of Deep Reinforcement Learning by stabilizing learning and allowing for higher training throughputs. We propose Gossi…
Separating value functions across time-scales
Joshua Romoff, Peter Henderson, Ahmed Touati +3
In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to…
Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods
Peter Henderson, Joshua Romoff, Joelle Pineau
Recent analyses of certain gradient descent optimization methods have shown that performance can degrade in some settings - such as with stochasticity or implicit momentum. In deep…
TarMAC: Targeted Multi-Agent Communication
Abhishek Das, Théophile Gervet, Joshua Romoff +4
We propose a targeted communication architecture for multi-agent reinforcement learning, where agents learn both what messages to send and whom to address them to while performing…