11 papers
Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization
Daojie Peng, Fulong Ma, Bingtao Wang +2
Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop cont…
SSTG-Nav: Metric-Grounded Spatial-Semantic Topological Graphs for Reusable Object Navigation
Daojie Peng, Bingtao Wang, Jun Ma
Service robots operating for months in the same homes, offices, and facilities should become more reliable with experience instead of searching familiar space from scratch for ever…
LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
Pei Liu, Nan Zheng, Lang Zhang +8
World Action Models (WAMs) have emerged as a powerful paradigm for embodied intelligence, yet the prevailing reliance on pixel-level video generation creates a fundamental bottlene…
GeoSem-WAM: Geometry- and Semantic-Aware World Action Models
Fulong Ma, Daojie Peng, Wenjun Yue +4
Recent World Action Models (WAMs) have demonstrated impressive capabilities in embodied decision-making. However, whether their effectiveness stems from explicit future imagination…
AttenA+: Rectifying Action Inequality in Robotic Foundation Models
Daojie Peng, Fulong Ma, Jiahang Cao +7
Existing robotic foundation models, while powerful, are predicated on an implicit assumption of temporal homogeneity: treating all actions as equally informative during optimizatio…
IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road Segmentation
Bingtao Wang, Daojie Peng, Fulong Ma +2
Semantic road segmentation is important for autonomous driving, but existing methods suffer severe performance degradation under low-light conditions. Many existing multi-modal fus…