collaborators

5 papers

cs.CV2025

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models

Hugues Thomas, Chen Chen, Jian Zhang

Effectively representing 3D scenes for Multimodal Large Language Models (MLLMs) is crucial yet challenging. Existing approaches commonly only rely on 2D image features and use vari…

cs.RO2025

DR-MPC: Deep Residual Model Predictive Control for Real-world Social Navigation

James R. Han, Hugues Thomas, Jian Zhang +2

How can a robot safely navigate around people with complex motion patterns? Deep Reinforcement Learning (DRL) in simulation holds some promise, but much prior work relies on simula…

cs.RO2024

Embedding Pose Graph, Enabling 3D Foundation Model Capabilities with a Compact Representation

Hugues Thomas, Mouli Sivapurapu, Jian Zhang

This paper presents the Embedding Pose Graph (EPG), an innovative method that combines the strengths of foundation models with a simple 3D representation suitable for robotics appl…

cs.RO2024

MakeWay: Object-Aware Costmaps for Proactive Indoor Navigation Using LiDAR

Binbin Xu, Allen Tao, Hugues Thomas +2

In this paper, we introduce a LiDAR-based robot navigation system, based on novel object-aware affordance-based costmaps. Utilizing a 3D object detection network, our system identi…

cs.CV2024

KPConvX: Modernizing Kernel Point Convolution with Kernel Attention

Hugues Thomas, Yao-Hung Hubert Tsai, Timothy D. Barfoot +1

In the field of deep point cloud understanding, KPConv is a unique architecture that uses kernel points to locate convolutional weights in space, instead of relying on Multi-Layer…