activity
20232025
most citedRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

273 citations · 375 across the 12 of their papers we have counts for

collaborators

12 papers

cs.RO20251 cited

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Gemini Robotics Team, Abbas Abdolmaleki, Saminda Abeyruwan +169

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the G…

cs.RO20256 cited

Gemini Robotics: Bringing AI into the Physical World

Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie +115

Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as…

cs.RO2024

STEER: Flexible Robotic Manipulation via Dense Language Grounding

Laura Smith, Alex Irpan, Montserrat Gonzalez Arenas +4

The complexity of the real world demands robotic systems that can intelligently adapt to unseen situations. We present STEER, a robot learning framework that bridges high-level, co…

cs.RO20241 cited

VADER: Visual Affordance Detection and Error Recovery for Multi Robot Human Collaboration

Michael Ahn, Montserrat Gonzalez Arenas, Matthew Bennice +22

Robots today can exploit the rich world knowledge of large language models to chain simple behavioral skills into long-horizon tasks. However, robots often get interrupted during l…

cs.RO20241 cited

Learning to Learn Faster from Human Feedback with Language Model Predictive Control

Jacky Liang, Fei Xia, Wenhao Yu +47

Large language models (LLMs) have been shown to exhibit a wide range of capabilities, such as writing robot code from language commands -- enabling non-experts to direct robot beha…

cs.RO202414 cited

AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents

Michael Ahn, Debidatta Dwibedi, Chelsea Finn +24

Foundation models that incorporate language, vision, and more recently actions have revolutionized the ability to harness internet scale data to reason about useful tasks. However,…