7 papers
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
Paul Pacaud, Ricardo Garcia, Shizhe Chen +1
Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in robot failure detection, their generaliz…
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
Thomas Chabal, Shizhe Chen, Jean Ponce +1
This paper addresses the Object Goal Navigation problem, where a robot must efficiently find a target object in an unknown environment. Existing implicit memory-based methods strug…
HORT: Monocular Hand-held Objects Reconstruction with Transformers
Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +1
Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which…
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
Shizhe Chen, Ricardo Garcia, Paul Pacaud +1
Robotic manipulation faces a significant challenge in generalizing across unseen objects, environments and tasks specified by diverse language instructions. To improve generalizati…
ComposeAnything: Composite Object Priors for Text-to-Image Generation
Zeeshan Khan, Shizhe Chen, Cordelia Schmid
Generating images from text involving complex and novel object arrangements remains a significant challenge for current text-to-image (T2I) models. Although prior layout-based meth…
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
Ricardo Garcia, Shizhe Chen, Cordelia Schmid
Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address t…