26 citations · 141 across the 42 of their papers we have counts for
7 papers · 1 filter
PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation
Shizhe Chen, Ricardo Garcia, Cordelia Schmid +1
The ability for robots to comprehend and execute manipulation tasks based on natural language instructions is a long-term goal in robotics. The dominant approaches for language-gui…
Explore and Tell: Embodied Visual Captioning in 3D Environments
Anwen Hu, Shizhe Chen, Liang Zhang +1
While current visual captioning models have achieved impressive performance, they often assume that the image is well-captured and provides a complete view of the scene. In real-wo…
Object Goal Navigation with Recursive Implicit Maps
Shizhe Chen, Thomas Chabal, Ivan Laptev +1
Object goal navigation aims to navigate an agent to locations of a given object category in unseen environments. Classical methods explicitly build maps of environments and require…
Robust Visual Sim-to-Real Transfer for Robotic Manipulation
Ricardo Garcia, Robin Strudel, Shizhe Chen +3
Learning visuomotor policies in simulation is much safer and cheaper than in the real world. However, due to discrepancies between the simulated and real data, simulator-trained po…
InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation
Anwen Hu, Shizhe Chen, Liang Zhang +1
Automatic image captioning evaluation is critical for benchmarking and promoting advances in image captioning research. Existing metrics only provide a single score to measure capt…
gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object Reconstruction
Zerui Chen, Shizhe Chen, Cordelia Schmid +1
Signed distance functions (SDFs) is an attractive framework that has recently shown promising results for 3D shape reconstruction from images. SDFs seamlessly generalize to differe…