28 citations · 54 across the 21 of their papers we have counts for
14 papers · 1 filter
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
Haifeng Huang, Yilun Chen, Zehan Wang +2
Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object…
Enhancing Indoor Occupancy Prediction via Sparse Query-Based Multi-Level Consistent Knowledge Distillation
Xiang Li, Yupeng Zheng, Pengfei Li +3
Occupancy prediction provides critical geometric and semantic understanding for robotics but faces efficiency-accuracy trade-offs. Current dense methods suffer computational waste…
MM-ACT: Learn from Multimodal Parallel Generation to Act
Haotian Liang, Xinyi Chen, Bin Wang +12
A generalist robotic policy needs both semantic understanding for task planning and the ability to interact with the environment through predictive capabilities. To tackle this, we…
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
Xin Wen, Bingchen Zhao, Yilun Chen +2
Pre-trained vision models (PVMs) are fundamental to modern robotics, yet their optimal configuration remains unclear. Through systematic evaluation, we find that while DINO and iBO…
CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object Query
Zhe Wang, Shaocong Xu, Xucai Zhuang +5
Cooperative perception enhances the individual perception capabilities of autonomous vehicles (AVs) by providing a comprehensive view of the environment. However, balancing percept…
Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
Xiang Li, Pengfei Li, Yupeng Zheng +3
Understanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surr…