6 citations · 7 across the 2 of their papers we have counts for
3 papers
cs.RO2025★ 1 cited
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
Chen Wang, Fei Xia, Wenhao Yu +6
Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during ta…
cs.RO2025★ 6 cited
Gemini Robotics: Bringing AI into the Physical World
Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie +115
Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as…
cs.RO2024
IPPON: Common Sense Guided Informative Path Planning for Object Goal Navigation
Kaixian Qu, Jie Tan, Tingnan Zhang +3
Navigating efficiently to an object in an unexplored environment is a critical skill for general-purpose intelligent robots. Recent approaches to this object goal navigation proble…