activity
20162022
most citedVisualizing Dynamics: from t-SNE to SEMI-MDPs

12 citations · 27 across the 7 of their papers we have counts for

collaborators

22 papers

cs.LG20221 cited

Palm up: Playing in the Latent Manifold for Unsupervised Pretraining

Hao Liu, Tom Zahavy, Volodymyr Mnih +1

Large and diverse datasets have been the cornerstones of many impressive advancements in artificial intelligence. Intelligent creatures, however, learn by interacting with the envi…

cs.LG2022

Meta-Gradients in Non-Stationary Environments

Jelena Luketina, Sebastian Flennerhag, Yannick Schroecker +3

Meta-gradient methods (Xu et al., 2018; Zahavy et al., 2020) offer a promising solution to the problem of hyperparameter selection and adaptation in non-stationary reinforcement le…

cs.LG20213 cited

Emphatic Algorithms for Deep Reinforcement Learning

Ray Jiang, Tom Zahavy, Zhongwen Xu +4

Off-policy learning allows us to learn about possible policies of behavior from experience generated by a different behavior policy. Temporal difference (TD) learning algorithms ca…

cs.LG20215 cited

Discovery of Options via Meta-Learned Subgoals

Vivek Veeriah, Tom Zahavy, Matteo Hessel +6

Temporal abstractions in the form of options have been shown to help reinforcement learning (RL) agents learn faster. However, despite prior work on this topic, the problem of disc…

cs.LG2021

Online Limited Memory Neural-Linear Bandits with Likelihood Matching

Ofir Nabati, Tom Zahavy, Shie Mannor

We study neural-linear bandits for solving problems where {\em both} exploration and representation learning play an important role. Neural-linear bandits harnesses the representat…

cs.LG20206 cited

Balancing Constraints and Rewards with Meta-Gradient D4PG

Dan A. Calian, Daniel J. Mankowitz, Tom Zahavy +4

Deploying Reinforcement Learning (RL) agents to solve real-world applications often requires satisfying complex system constraints. Often the constraint thresholds are incorrectly…