2 citations · 4 across the 7 of their papers we have counts for
7 papers
OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving
Julong Wei, Shanshuai Yuan, Pengfei Li +3
The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping fro…
GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping
Yuhang Zheng, Xiangyu Chen, Yupeng Zheng +12
Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit, particularly within the domain of robotics. Such technology facilitates robots i…
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
Yupeng Zheng, Xiang Li, Pengfei Li +6
Monocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, partic…
3D Implicit Transporter for Temporally Consistent Keypoint Discovery
Chengliang Zhong, Yuhang Zheng, Yupeng Zheng +9
Keypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric…
LODE: Locally Conditioned Eikonal Implicit Scene Completion from Sparse LiDAR
Pengfei Li, Ruowen Zhao, Yongliang Shi +4
Scene completion refers to obtaining dense scene representation from an incomplete perception of complex 3D scenes. This helps robots detect multi-scale obstacles and analyse objec…
STEPS: Joint Self-supervised Nighttime Image Enhancement and Depth Estimation
Yupeng Zheng, Chengliang Zhong, Pengfei Li +8
Self-supervised depth estimation draws a lot of attention recently as it can promote the 3D sensing capabilities of self-driving vehicles. However, it intrinsically relies upon the…