12 citations · 27 across the 7 of their papers we have counts for
22 papers
Palm up: Playing in the Latent Manifold for Unsupervised Pretraining
Hao Liu, Tom Zahavy, Volodymyr Mnih +1
Large and diverse datasets have been the cornerstones of many impressive advancements in artificial intelligence. Intelligent creatures, however, learn by interacting with the envi…
Meta-Gradients in Non-Stationary Environments
Jelena Luketina, Sebastian Flennerhag, Yannick Schroecker +3
Meta-gradient methods (Xu et al., 2018; Zahavy et al., 2020) offer a promising solution to the problem of hyperparameter selection and adaptation in non-stationary reinforcement le…
Emphatic Algorithms for Deep Reinforcement Learning
Ray Jiang, Tom Zahavy, Zhongwen Xu +4
Off-policy learning allows us to learn about possible policies of behavior from experience generated by a different behavior policy. Temporal difference (TD) learning algorithms ca…
Discovery of Options via Meta-Learned Subgoals
Vivek Veeriah, Tom Zahavy, Matteo Hessel +6
Temporal abstractions in the form of options have been shown to help reinforcement learning (RL) agents learn faster. However, despite prior work on this topic, the problem of disc…
Online Limited Memory Neural-Linear Bandits with Likelihood Matching
Ofir Nabati, Tom Zahavy, Shie Mannor
We study neural-linear bandits for solving problems where {\em both} exploration and representation learning play an important role. Neural-linear bandits harnesses the representat…
Balancing Constraints and Rewards with Meta-Gradient D4PG
Dan A. Calian, Daniel J. Mankowitz, Tom Zahavy +4
Deploying Reinforcement Learning (RL) agents to solve real-world applications often requires satisfying complex system constraints. Often the constraint thresholds are incorrectly…