588 citations
- Google DeepMind (United Kingdom)GB17 papers
- Google (United States)US10 papers
- Stanford UniversityUS3 papers
- Carnegie Mellon UniversityUS2 papers
- ETH ZurichCH2 papers
- Google (Switzerland)CH2 papers
- King's College LondonGB2 papers
- Massachusetts Institute of TechnologyUS2 papers
- University College LondonGB2 papers
- University of California San DiegoUS2 papers
- University of GlasgowGB2 papers
- University of TorontoCA2 papers
9 papers · 1 filter
Half-Hop: A graph upsampling approach for slowing down message passing
Mehdi Azabou, Venkataramana Ganesh, Shantanu Thakoor +6
Message passing neural networks have shown a lot of success on graph-structured data. However, there are many instances where message passing can lead to over-smoothing or fail whe…
DoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm
Yunhao Tang, Tadashi Kozuno, Mark Rowland +4
Multi-step learning applies lookahead over multiple time steps and has proved valuable in policy evaluation settings. However, in the optimal control case, the impact of multi-step…
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice
Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang +12
Mirror descent value iteration (MDVI), an abstraction of Kullback-Leibler (KL) and entropy-regularized reinforcement learning (RL), has served as the basis for recent high-performi…
Understanding Self-Predictive Learning for Reinforcement Learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond +13
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their…
Marginalized Operators for Off-policy Reinforcement Learning
Yunhao Tang, Mark Rowland, Rémi Munos +1
In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi…
Retrieval-Augmented Reinforcement Learning
Anirudh Goyal, Abram L. Friesen, Andrea Banino +13
Most deep reinforcement learning (RL) algorithms distill experience into parametric behavior policies or value functions via gradient updates. While effective, this approach has se…