1 paper
Zichun Xu, Jingdong Zhao, Chenyu Guo +6
Depth information is robust to scene appearance variations and inherently carries 3D spatial details. Thus, a visual backbone based on the vision transformer is proposed to fuse RG…