activity
20172026
most citedLearning model-based planning from scratch

78 citations · 182 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

11 papers · 1 filter

cs.LG20241 cited

Amplifying human performance in combinatorial competitive programming

Petar Veličković, Alex Vitvitskyi, Larisa Markeeva +4

Recent years have seen a significant surge in complex AI systems for competitive programming, capable of performing at admirable levels against human competitors. While steady prog…

cs.LG2020

Representation Learning via Invariant Causal Mechanisms

Jovana Mitrovic, Brian McWilliams, Jacob Walker +2

Self-supervised learning has emerged as a strategy to reduce the reliance on costly supervised signal by pretraining representations only using unlabeled data. These methods combin…

cs.LG20202 cited

Beyond Tabula-Rasa: a Modular Reinforcement Learning Approach for Physically Embedded 3D Sokoban

Peter Karkus, Mehdi Mirza, Arthur Guez +5

Intelligent robots need to achieve abstract objectives using concrete, spatiotemporally complex sensory information and motor control. Tabula rasa deep reinforcement learning (RL)…

cs.LG20209 cited

Divide-and-Conquer Monte Carlo Tree Search For Goal-Directed Planning

Giambattista Parascandolo, Lars Buesing, Josh Merel +6

Standard planners for sequential decision making (including Monte Carlo planning, tree search, dynamic programming, etc.) are constrained by an implicit sequential planning assumpt…

cs.LG202016 cited

Causally Correct Partial Models for Reinforcement Learning

Danilo J. Rezende, Ivo Danihelka, George Papamakarios +11

In reinforcement learning, we can learn a model of future observations and rewards, and use it to plan the agent's next actions. However, jointly modeling future observations can b…

cs.LG2020

Value-driven Hindsight Modelling

Arthur Guez, Fabio Viola, Théophane Weber +5

Value estimation is a critical component of the reinforcement learning (RL) paradigm. The question of how to effectively learn value predictors from data is one of the major proble…