TransCG: A Large-Scale Real-World Dataset for Transparent Object Depth Completion and a Grasping Baseline
arXiv:2202.08471 · doi:10.1109/LRA.2022.3183256
Abstract
Transparent objects are common in our daily life and frequently handled in the automated production line. Robust vision-based robotic grasping and manipulation for these objects would be beneficial for automation. However, the majority of current grasping algorithms would fail in this case since they heavily rely on the depth image, while ordinary depth sensors usually fail to produce accurate depth information for transparent objects owing to the reflection and refraction of light. In this work, we address this issue by contributing a large-scale real-world dataset for transparent object depth completion, which contains 57,715 RGB-D images from 130 different scenes. Our dataset is the first large-scale, real-world dataset that provides ground truth depth, surface normals, transparent masks in diverse and cluttered scenes. Cross-domain experiments show that our dataset is more general and can enable better generalization ability for models. Moreover, we propose an end-to-end depth completion network, which takes the RGB image and the inaccurate depth map as inputs and outputs a refined depth map. Experiments demonstrate superior efficacy, efficiency and robustness of our method over previous works, and it is able to process images of high resolutions under limited hardware resources. Real robot experiments show that our method can also be applied to novel transparent object grasping robustly. The full dataset and our method are publicly available at www.graspnet.net/transcg
project page: www.graspnet.net/transcg
References in corpus (4)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Dex-NeRF: Using a Neural Radiance Field to Grasp Transparent Objects
- S4G: Amodal Single-view Single-Shot SE(3) Grasp Detection in Cluttered Scenes
Cited by in corpus (7)
- Challenges for Monocular 6D Object Pose Estimation in Robotics
- FDCT: Fast Depth Completion for Transparent Objects
- ASGrasp: Generalizable Transparent Object Reconstruction and 6-DoF Grasp Detection from RGB-D Active Stereo Camera
- A Survey of Embodied Learning for Object-Centric Robotic Manipulation
- Diffusion-Based Depth Inpainting for Transparent and Reflective Objects
- RFTrans: Leveraging Refractive Flow of Transparent Objects for Surface Normal Estimation and Manipulation
- Adaptive Grasping of Moving Objects in Dense Clutter via Global-to-Local Detection and Static-to-Dynamic Planning