activity
20152021
most citedCounterfactual Risk Minimization: Learning from Logged Bandit Feedback

127 citations · 237 across the 8 of their papers we have counts for

collaborators

11 papers

cs.LG20212 cited

Improving Long-Term Metrics in Recommendation Systems using Short-Horizon Reinforcement Learning

Bogdan Mazoure, Paul Mineiro, Pavithra Srinath +3

We study session-based recommendation scenarios where we want to recommend items to users during sequential interactions to improve their long-term utility. Optimizing a long-term…

cs.LG202035 cited

Provably Good Batch Reinforcement Learning Without Great Exploration

Yao Liu, Adith Swaminathan, Alekh Agarwal +1

Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is chall…

cs.LG2019

Working Memory Graphs

Ricky Loynd, Roland Fernandez, Asli Celikyilmaz +2

Transformers have increasingly outperformed gated RNNs in obtaining new state-of-the-art results on supervised tasks involving text sequences. Inspired by this trend, we study the…

cs.LG2019

Learning Calibratable Policies using Programmatic Style-Consistency

Eric Zhan, Albert Tseng, Yisong Yue +2

We study the problem of controllable generation of long-term sequential behaviors, where the goal is to calibrate to multiple behavior styles simultaneously. In contrast to the wel…

cs.LG2019

Metareasoning in Modular Software Systems: On-the-Fly Configuration using Reinforcement Learning with Rich Contextual Representations

Aditya Modi, Debadeepta Dey, Alekh Agarwal +4

Assemblies of modular subsystems are being pressed into service to perform sensing, reasoning, and decision making in high-stakes, time-critical tasks in such areas as transportati…

cs.LG2019

Multi-Preference Actor Critic

Ishan Durugkar, Matthew Hausknecht, Adith Swaminathan +1

Policy gradient algorithms typically combine discounted future rewards with an estimated value function, to compute the direction and magnitude of parameter updates. However, for m…