2 papers
cs.CV2026
PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning
Dongxu Zhang, Yiding Sun, Pengcheng Li +12
While Multimodal Large Language Models (MLLMs) demonstrate proficiency in 2D scenes, extending their perceptual intelligence to 3D point cloud understanding remains a significant c…
cs.CV2025
Depth3DLane: Monocular 3D Lane Detection via Depth Prior Distillation
Dongxin Lyu, Han Huang, Cheng Tan +1
Monocular 3D lane detection is challenging due to the difficulty in capturing depth information from single-camera images. A common strategy involves transforming front-view (FV) i…