4 papers
Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation
Yunpeng Gao, Zhigang Wang, Pengfei Han +3
Aerial Vision-and-Language Navigation (VLN) is a novel task enabling Unmanned Aerial Vehicles (UAVs) to navigate in outdoor environments through natural language instructions and v…
COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models
Kehui Liu, Zixin Tang, Dong Wang +3
Leveraging the powerful reasoning capabilities of large language models (LLMs), recent LLM-based robot task planning methods yield promising results. However, they mainly focus on…
AlignBot: Aligning VLM-powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots
Zhaxizhuoma Zhaxizhuoma, Pengan Chen, Ziniu Wu +7
This paper presents AlignBot, a novel framework designed to optimize VLM-powered customized task planning for household robots by effectively aligning with user reminders. In domes…
LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and Control
Delin Qu, Qizhi Chen, Pingrui Zhang +5
This paper scales object-level reconstruction to complex scenes, advancing interactive scene reconstruction. We introduce two datasets, OmniSim and InterReal, featuring 28 scenes w…