3 papers
cs.RO2025
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
Fuhao Li, Wenxuan Song, Han Zhao +5
Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are buil…
cs.CV2025
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
Fuhao Li, Huan Jin, Bin Gao +3
Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing dat…
cs.RO2024
THUD++: Large-Scale Dynamic Indoor Scene Dataset and Benchmark for Mobile Robots
Zeshun Li, Fuhao Li, Wanting Zhang +4
Most existing mobile robotic datasets primarily capture static scenes, limiting their utility for evaluating robotic performance in dynamic environments. To address this, we presen…