Publications (9)
Grounded Language Learning Fast and Slow
Felix Hill, Olivier Tieleman, Tamara von Glehn +3
Recent work has shown that large text-based neural language models, trained with conventional supervised learning objectives, acquire a surprising propensity for few- and one-shot…
Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual Understanding
Talfan Evans, Shreya Pathak, Hamza Merzic +3
Power-law scaling indicates that large-scale training with uniform sampling is prohibitively slow. Active learning methods aim to increase data efficiency by prioritizing learning…
Causally Correct Partial Models for Reinforcement Learning
Danilo J. Rezende, Ivo Danihelka, George Papamakarios +11
In reinforcement learning, we can learn a model of future observations and rewards, and use it to plan the agent's next actions. However, jointly modeling future observations can b…
Scaling Instructable Agents Across Many Simulated Worlds
SIMA Team, Maria Abi Raad, Arun Ahuja +91
Building embodied AI systems that can follow arbitrary language instructions in any 3D environment is a key challenge for creating general AI. Accomplishing this goal requires lear…
Creating Multimodal Interactive Agents with Imitation and Self-Supervised Learning
DeepMind Interactive Agents Team, Josh Abramson, Arun Ahuja +22
A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through…
Shaping Belief States with Generative Environment Models for RL
Karol Gregor, Danilo Jimenez Rezende, Frederic Besse +3
When agents interact with a complex environment, they must form and maintain beliefs about the relevant aspects of that environment. We propose a way to efficiently train expressiv…
Probing Emergent Semantics in Predictive Agents via Question Answering
Abhishek Das, Federico Carnevale, Hamza Merzic +8
Recent work has shown how predictive modeling can endow agents with rich knowledge of their surroundings, improving their ability to act in complex environments. We propose questio…
Data curation via joint example selection further accelerates multimodal learning
Talfan Evans, Nikhil Parthasarathy, Hamza Merzic +1
Data curation is an essential component of large-scale pretraining. In this work, we demonstrate that jointly selecting batches of data is more effective for learning than selectin…
Leveraging Contact Forces for Learning to Grasp
Hamza Merzic, Miroslav Bogdanovic, Daniel Kappler +2
Grasping objects under uncertainty remains an open problem in robotics research. This uncertainty is often due to noisy or partial observations of the object pose or shape. To enab…