activity
20182020
most citedWay Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

134 citations · 134 across the 2 of their papers we have counts for

collaborators

8 papers

cs.CL2020

Human-centric Dialog Training via Offline Reinforcement Learning

Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun +5

How can we train a dialog model to produce better conversations by learning from human feedback, without the risk of humans teaching it harmful chat behaviors? We start by hosting…

cs.LG2019

Characterizing Sources of Uncertainty to Proxy Calibration and Disambiguate Annotator and Data Bias

Asma Ghandeharioun, Brian Eoff, Brendan Jou +1

Supporting model interpretability for complex phenomena where annotators can legitimately disagree, such as emotion recognition, is a challenging machine learning task. In this wor…

cs.LG2019

Hierarchical Reinforcement Learning for Open-Domain Dialog

Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun +2

Open-domain dialog generation is a challenging problem; maximum likelihood training can lead to repetitive outputs, models have difficulty tracking long-term conversational goals,…

cs.HC2019

Towards Understanding Emotional Intelligence for Behavior Change Chatbots

Asma Ghandeharioun, Daniel McDuff, Mary Czerwinski +1

A natural conversational interface that allows longitudinal symptom tracking would be extremely valuable in health/wellness applications. However, the task of designing emotionally…

cs.HC2019

Engineering Music to Slow Breathing and Invite Relaxed Physiology

Grace Leslie, Asma Ghandeharioun, Diane Y. Zhou +1

We engineered an interactive music system that influences a user's breathing rate to induce a relaxation response. This system generates ambient music containing periodic shifts in…

cs.LG2019134 cited

Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen +5

Most deep reinforcement learning (RL) systems are not able to learn effectively from off-policy data, especially if they cannot explore online in the environment. These are critica…