127 citations · 237 across the 8 of their papers we have counts for
11 papers
Improving Long-Term Metrics in Recommendation Systems using Short-Horizon Reinforcement Learning
Bogdan Mazoure, Paul Mineiro, Pavithra Srinath +3
We study session-based recommendation scenarios where we want to recommend items to users during sequential interactions to improve their long-term utility. Optimizing a long-term…
Provably Good Batch Reinforcement Learning Without Great Exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal +1
Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is chall…
Working Memory Graphs
Ricky Loynd, Roland Fernandez, Asli Celikyilmaz +2
Transformers have increasingly outperformed gated RNNs in obtaining new state-of-the-art results on supervised tasks involving text sequences. Inspired by this trend, we study the…
Learning Calibratable Policies using Programmatic Style-Consistency
Eric Zhan, Albert Tseng, Yisong Yue +2
We study the problem of controllable generation of long-term sequential behaviors, where the goal is to calibrate to multiple behavior styles simultaneously. In contrast to the wel…
Metareasoning in Modular Software Systems: On-the-Fly Configuration using Reinforcement Learning with Rich Contextual Representations
Aditya Modi, Debadeepta Dey, Alekh Agarwal +4
Assemblies of modular subsystems are being pressed into service to perform sensing, reasoning, and decision making in high-stakes, time-critical tasks in such areas as transportati…
Multi-Preference Actor Critic
Ishan Durugkar, Matthew Hausknecht, Adith Swaminathan +1
Policy gradient algorithms typically combine discounted future rewards with an estimated value function, to compute the direction and magnitude of parameter updates. However, for m…