6 papers
StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
Zhe Liu, Jinghua Hou, Yuxiang Lu +7
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting…
SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control
Ruihua Han, Rui Gao, Zhe Liu +6
Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, s…
SSC: A Verifiable Structured Representation for Bimanual Manipulation Labelling
Yupu Lu, Shuang Wu, Sihan Chen +4
Subtask labels decompose a long-horizon manipulation demonstration into shorter semantic segments for policy training and evaluation. Natural language descriptions are easy to read…
IR-SIM: A Lightweight Skill-Native Simulator for Navigation, Learning, and Benchmarking
Ruihua Han, Shuai Wang, Chengyang Li +8
Simulation plays a key role in automated robotics research supported by large language models (LLMs). However, existing simulators often require custom code or complex interfaces,…
Agentic Self-Evolutionary Replanning for Embodied Navigation
Guoliang Li, Ruihua Han, Chengyang Li +5
Failure is inevitable for embodied navigation in complex environments. To enhance the resilience, replanning (RP) is a viable option, where the robot is allowed to fail, but is cap…
A Haptic-Based Proximity Sensing System for Buried Object in Granular Material
Zeqing Zhang, Ruixing Jia, Youcan Yan +5
The proximity perception of objects in granular materials is significant, especially for applications like minesweeping. However, due to particles' opacity and complex properties,…