activity
20172019
most citedChanging Model Behavior at Test-Time Using Reinforcement Learning

18 citations · 31 across the 4 of their papers we have counts for

collaborators

6 papers

cs.LG20192 cited

Energy-Inspired Models: Learning with Sampler-Induced Distributions

Dieterich Lawson, George Tucker, Bo Dai +1

Energy-based models (EBMs) are powerful probabilistic models, but suffer from intractable sampling and density evaluation due to the partition function. As a result, inference in E…

cs.LG2018

Doubly Reparameterized Gradient Estimators for Monte Carlo Objectives

George Tucker, Dieterich Lawson, Shixiang Gu +1

Deep latent variable models have become a popular model choice due to the scalable learning algorithms introduced by (Kingma & Welling, 2013; Rezende et al., 2014). These approache…

cs.CL20173 cited

An online sequence-to-sequence model for noisy speech recognition

Chung-Cheng Chiu, Dieterich Lawson, Yuping Luo +4

Generative models have long been the dominant approach for speech recognition. The success of these models however relies on the use of sophisticated recipes and complicated machin…

cs.AI2017

Learning Hard Alignments with Variational Inference

Dieterich Lawson, Chung-Cheng Chiu, George Tucker +3

There has recently been significant interest in hard attention models for tasks such as object recognition, visual captioning and speech recognition. Hard attention can offer benef…

cs.LG20178 cited

Particle Value Functions

Chris J. Maddison, Dieterich Lawson, George Tucker +4

The policy gradients of the expected return objective can react slowly to rare rewards. Yet, in some cases agents may wish to emphasize the low or high returns regardless of their…

stat.ML201718 cited

Changing Model Behavior at Test-Time Using Reinforcement Learning

Augustus Odena, Dieterich Lawson, Christopher Olah

Machine learning models are often used at test-time subject to constraints and trade-offs not present at training-time. For example, a computer vision model operating on an embedde…