activity
20122019
most citedProceedings of the 29th International Conference on Machine Learning (ICML-12)

1.6k citations · 2k across the 10 of their papers we have counts for

collaborators

13 papers

cs.LG20192 cited

Recurrent Value Functions

Pierre Thodoroff, Nishanth Anand, Lucas Caccia +2

Despite recent successes in Reinforcement Learning, value-based methods often suffer from high variance hindering performance. In this paper, we illustrate this in a continuous con…

cs.CL201814 cited

A Deep Reinforcement Learning Chatbot (Short Version)

Iulian V. Serban, Chinnadhurai Sankar, Mathieu Germain +15

We present MILABOT: a deep reinforcement learning chatbot developed by the Montreal Institute for Learning Algorithms (MILA) for the Amazon Alexa Prize competition. MILABOT is capa…

cs.CL2017

Ethical Challenges in Data-Driven Dialogue Systems

Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier +4

The use of dialogue systems as a medium for human-machine interaction is an increasingly prevalent paradigm. A growing number of dialogue systems use conversation strategies that a…

stat.ML20178 cited

ACtuAL: Actor-Critic Under Adversarial Learning

Anirudh Goyal, Nan Rosemary Ke, Alex Lamb +4

Generative Adversarial Networks (GANs) are a powerful framework for deep generative modeling. Posed as a two-player minimax problem, GANs are typically trained end-to-end on real-v…

cs.LG2017

OptionGAN: Learning Joint Reward-Policy Options using Generative Adversarial Inverse Reinforcement Learning

Peter Henderson, Wei-Di Chang, Pierre-Luc Bacon +3

Reinforcement learning has shown promise in learning policies that can solve complex problems. However, manually specifying a good reward function can be difficult, especially for…

cs.CL2017200 cited

A Deep Reinforcement Learning Chatbot

Iulian V. Serban, Chinnadhurai Sankar, Mathieu Germain +15

We present MILABOT: a deep reinforcement learning chatbot developed by the Montreal Institute for Learning Algorithms (MILA) for the Amazon Alexa Prize competition. MILABOT is capa…