activity
20182022
most citedSolving math word problems with process- and outcome-based feedback

23 citations · 65 across the 6 of their papers we have counts for

collaborators

9 papers

cs.LG202223 cited

Solving math word problems with process- and outcome-based feedback

Jonathan Uesato, Nate Kushman, Ramana Kumar +6

Recent work has shown that asking language models to generate reasoning steps improves performance on many reasoning tasks. When moving beyond prompting, this raises the question o…

cs.RO202220 cited

Imitate and Repurpose: Learning Reusable Robot Movement Skills From Human and Animal Behaviors

Steven Bohez, Saran Tunyasuvunakool, Philemon Brakel +18

We investigate the use of prior knowledge of human and animal movement to learn reusable locomotion skills for real legged robots. Our approach builds upon previous work on imitati…

cs.AI202112 cited

From Motor Control to Team Play in Simulated Humanoid Football

Siqi Liu, Guy Lever, Zhe Wang +19

Intelligent behaviour in the physical world exhibits structure at multiple spatial and temporal scales. Although movements are ultimately executed at the level of instantaneous mus…

cs.RO20201 cited

"What, not how": Solving an under-actuated insertion task from scratch

Giulia Vezzani, Michael Neunert, Markus Wulfmeier +7

Robot manipulation requires a complex set of skills that need to be carefully combined and coordinated to solve a task. Yet, most ReinforcementLearning (RL) approaches in robotics…

cs.LG20203 cited

Simple Sensor Intentions for Exploration

Tim Hertweck, Martin Riedmiller, Michael Bloesch +5

Modern reinforcement learning algorithms can learn solutions to increasingly difficult control problems while at the same time reduce the amount of prior knowledge needed for their…

cs.LG2020

Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning

Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp +6

Off-policy reinforcement learning algorithms promise to be applicable in settings where only a fixed data-set (batch) of environment interactions is available and no new experience…