12 citations · 16 across the 3 of their papers we have counts for
3 papers
cs.LG2017★ 3 cited
A short variational proof of equivalence between policy gradients and soft Q learning
Pierre H. Richemond, Brendan Maginnis
Two main families of reinforcement learning algorithms, Q-learning and policy gradients, have recently been proven to be equivalent when using a softmax relaxation on one part, and…
cs.LG2017★ 12 cited
On Wasserstein Reinforcement Learning and the Fokker-Planck equation
Pierre H. Richemond, Brendan Maginnis
Policy gradients methods often achieve better performance when the change in policy is limited to a small Kullback-Leibler divergence. We derive policy gradients where the change i…
cs.LG2017★ 1 cited
Efficiently applying attention to sequential data with the Recurrent Discounted Attention unit
Brendan Maginnis, Pierre H. Richemond
Recurrent Neural Networks architectures excel at processing sequences by modelling dependencies over different timescales. The recently introduced Recurrent Weighted Average (RWA)…