activity
20152020
most citedLyapunov-based Safe Policy Optimization for Continuous Control

154 citations · 473 across the 12 of their papers we have counts for

collaborators

24 papers

cs.LG20203 cited

Non-Stationary Latent Bandits

Joey Hong, Branislav Kveton, Manzil Zaheer +4

Users of recommender systems often behave in a non-stationary fashion, due to their evolving preferences and tastes over time. In this work, we propose a practical approach for fas…

cs.LG2020

CoinDICE: Off-Policy Confidence Interval Estimation

Bo Dai, Ofir Nachum, Yinlam Chow +3

We study high-confidence behavior-agnostic off-policy evaluation in reinforcement learning, where the goal is to estimate a confidence interval on a target policy's value, given on…

cs.CL2020

Safe Reinforcement Learning with Natural Language Constraints

Tsung-Yen Yang, Michael Hu, Yinlam Chow +2

While safe reinforcement learning (RL) holds great promise for many practical applications like robotics or autonomous cars, current approaches require specifying constraints in ma…

cs.LG20208 cited

Control-Aware Representations for Model-based Reinforcement Learning

Brandon Cui, Yinlam Chow, Mohammad Ghavamzadeh

A major challenge in modern reinforcement learning (RL) is efficient control of dynamical systems from high-dimensional sensory observations. Learning controllable embedding (LCE)…

cs.LG2020

Variational Model-based Policy Optimization

Yinlam Chow, Brandon Cui, MoonKyung Ryu +1

Model-based reinforcement learning (RL) algorithms allow us to combine model-generated data with those collected from interaction with the real system in order to alleviate the dat…

cs.LG202011 cited

Latent Bandits Revisited

Joey Hong, Branislav Kveton, Manzil Zaheer +3

A latent bandit problem is one in which the learning agent knows the arm reward distributions conditioned on an unknown discrete latent state. The primary goal of the agent is to i…