6 papers
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
Yixia Li, Hongru Wang, Peng Lai +13
Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet m…
IR-SIM: A Lightweight Skill-Native Simulator for Navigation, Learning, and Benchmarking
Ruihua Han, Shuai Wang, Chengyang Li +8
Simulation plays a key role in automated robotics research supported by large language models (LLMs). However, existing simulators often require custom code or complex interfaces,…
Semantic2D: Enabling Semantic Scene Understanding with 2D Lidar Alone
Zhanteng Xie, Yipeng Pan, Yinqiang Zhang +2
This article presents a complete semantic scene understanding workflow using only a single 2D lidar. This fills the gap in 2D lidar semantic segmentation, thereby enabling the reth…
LP-ICP: General Localizability-Aware Point Cloud Registration for Robust Localization in Extreme Unstructured Environments
Haosong Yue, Qingyuan Xu, Fei Chen +2
The Iterative Closest Point (ICP) algorithm is a crucial component of LiDAR-based SLAM algorithms. However, its performance can be negatively affected in unstructured environments…
iPad: Iterative Proposal-centric End-to-End Autonomous Driving
Ke Guo, Haochen Liu, Xiaojun Wu +2
End-to-end (E2E) autonomous driving systems offer a promising alternative to traditional modular pipelines by reducing information loss and error accumulation, with significant pot…
Aerial Vision-and-Language Navigation with Grid-based View Selection and Map Construction
Ganlong Zhao, Guanbin Li, Jia Pan +1
Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to groun…