From the 1 of 14 linked papers with an AI index.
5 papers · 1 filter
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
Shizhe Chen, Paul Pacaud, Cordelia Schmid
Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most exi…
Guardian: Detecting Robotic Planning and Execution Errors with Vision-Language Models
Paul Pacaud, Ricardo Garcia, Shizhe Chen +1
Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in robot failure detection, their generaliz…
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
Thomas Chabal, Shizhe Chen, Jean Ponce +1
This paper addresses the Object Goal Navigation problem, where a robot must efficiently find a target object in an unknown environment. Existing implicit memory-based methods strug…
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
Shizhe Chen, Ricardo Garcia, Paul Pacaud +1
Robotic manipulation faces a significant challenge in generalizing across unseen objects, environments and tasks specified by diverse language instructions. To improve generalizati…
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
Ricardo Garcia, Shizhe Chen, Cordelia Schmid
Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address t…