5 papers
Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models
Hugues Thomas, Chen Chen, Jian Zhang
Effectively representing 3D scenes for Multimodal Large Language Models (MLLMs) is crucial yet challenging. Existing approaches commonly only rely on 2D image features and use vari…
DR-MPC: Deep Residual Model Predictive Control for Real-world Social Navigation
James R. Han, Hugues Thomas, Jian Zhang +2
How can a robot safely navigate around people with complex motion patterns? Deep Reinforcement Learning (DRL) in simulation holds some promise, but much prior work relies on simula…
Embedding Pose Graph, Enabling 3D Foundation Model Capabilities with a Compact Representation
Hugues Thomas, Mouli Sivapurapu, Jian Zhang
This paper presents the Embedding Pose Graph (EPG), an innovative method that combines the strengths of foundation models with a simple 3D representation suitable for robotics appl…
MakeWay: Object-Aware Costmaps for Proactive Indoor Navigation Using LiDAR
Binbin Xu, Allen Tao, Hugues Thomas +2
In this paper, we introduce a LiDAR-based robot navigation system, based on novel object-aware affordance-based costmaps. Utilizing a 3D object detection network, our system identi…
KPConvX: Modernizing Kernel Point Convolution with Kernel Attention
Hugues Thomas, Yao-Hung Hubert Tsai, Timothy D. Barfoot +1
In the field of deep point cloud understanding, KPConv is a unique architecture that uses kernel points to locate convolutional weights in space, instead of relying on Multi-Layer…