26 citations · 141 across the 42 of their papers we have counts for
4 papers · 1 filter
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
Ricardo Garcia, Shizhe Chen, Cordelia Schmid
Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address t…
Think-Program-reCtify: 3D Situated Reasoning with Large Language Models
Qingrong He, Kejun Lin, Shizhe Chen +2
This work addresses the 3D situated reasoning task which aims to answer questions given egocentric observations in a 3D environment. The task remains challenging as it requires com…
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
Zerui Chen, Shizhe Chen, Etienne Arlaud +2
In this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits…
SUGAR: Pre-training 3D Visual Representations for Robotics
Shizhe Chen, Ricardo Garcia, Ivan Laptev +1
Learning generalizable visual representations from Internet data has yielded promising results for robotics. Yet, prevailing approaches focus on pre-training 2D representations, be…