9 papers
Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization
Daojie Peng, Fulong Ma, Bingtao Wang +2
Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop cont…
GeoSem-WAM: Geometry- and Semantic-Aware World Action Models
Fulong Ma, Daojie Peng, Wenjun Yue +4
Recent World Action Models (WAMs) have demonstrated impressive capabilities in embodied decision-making. However, whether their effectiveness stems from explicit future imagination…
AttenA+: Rectifying Action Inequality in Robotic Foundation Models
Daojie Peng, Fulong Ma, Jiahang Cao +7
Existing robotic foundation models, while powerful, are predicated on an implicit assumption of temporal homogeneity: treating all actions as equally informative during optimizatio…
IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road Segmentation
Bingtao Wang, Daojie Peng, Fulong Ma +2
Semantic road segmentation is important for autonomous driving, but existing methods suffer severe performance degradation under low-light conditions. Many existing multi-modal fus…
LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation
Daojie Peng, Bingtao Wang, Fulong Ma +2
Road segmentation is a fundamental perception task for autonomous driving and intelligent robotic systems, requiring both high accuracy and real-time inference, especially for depl…
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
Daojie Peng, Fulong Ma, Jun Ma
Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of vis…