Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang +4
Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world…
cs.CV2024
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
Chenyang Ma, Kai Lu, Ta-Ying Cheng +2
Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such…
cs.CV2023
Pre-Training LiDAR-Based 3D Object Detectors Through Colorization
Tai-Yu Pan, Chenyang Ma, Tianle Chen +7
Accurate 3D object detection and understanding for self-driving cars heavily relies on LiDAR point clouds, necessitating large amounts of labeled data to train. In this work, we in…