3 papers
cs.CV2025
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
Zechuan Li, Hongshan Yu, Yihao Ding +3
3D Scene Question Answering (3D SQA) represents an interdisciplinary task that integrates 3D visual perception and natural language processing, empowering intelligent agents to com…
cs.CV2025
Depth as Points: Center Point-based Depth Estimation
Zhiheng Tu, Xinjian Huang, Yong He +3
The perception of vehicles and pedestrians in urban scenarios is crucial for autonomous driving. This process typically involves complicated data collection, imposes high computati…
cs.CV2025
PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation
Yong He, Hongshan Yu, Mingtao Feng +5
Diffusion probabilistic models are traditionally used to generate colors at fixed pixel positions in 2D images. Building on this, we extend diffusion models to point cloud semantic…