5 papers
MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight
Zehua Fan, Junjie He, Wenxuan Song +14
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
Fuhao Li, Wenxuan Song, Han Zhao +5
Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are buil…
DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding
Qinghongbing Xie, Zijian Liang, Fuhao Li +1
Effective scene representation is critical for the visual grounding ability of representations, yet existing methods for 3D Visual Grounding are often constrained. They either only…
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
Fuhao Li, Huan Jin, Bin Gao +3
Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing dat…
THUD++: Large-Scale Dynamic Indoor Scene Dataset and Benchmark for Mobile Robots
Zeshun Li, Fuhao Li, Wanting Zhang +4
Most existing mobile robotic datasets primarily capture static scenes, limiting their utility for evaluating robotic performance in dynamic environments. To address this, we presen…