20 papers
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
Runze Li, Hongyin Zhang, Junxi Jin +5
Vision-Language-Action (VLA) models have emerged as a promising paradigm for building embodied agents that ground perception and language into action. However, most existing approa…
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
Minghui Lin, Pengxiang Ding, Shu Wang +7
Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property,…
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
Zirui Ge, Pengxiang Ding, Baohua Yin +16
Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…
VORL-EXPLORE: A Hybrid Learning Planning Approach to Multi-Robot Exploration in Dynamic Environments
Ning Liu, Sen Shen, Zheng Li +4
Hierarchical multi-robot exploration commonly decouples frontier allocation from local navigation, which can make the system brittle in dense and dynamic environments. Because the…
CMR: Contractive Mapping Embeddings for Robust Humanoid Locomotion on Unstructured Terrains
Qixin Zeng, Hongyin Zhang, Shangke Lyu +3
Robust disturbance rejection remains a longstanding challenge in humanoid locomotion, particularly on unstructured terrains where sensing is unreliable and model mismatch is pronou…
Activation-wise Propagation: A One-Timestep Strategy for Spiking Neural Networks
Jian Song, Xiangfei Yang, Shangke Lyu +1
Spiking neural networks (SNNs) have demonstrated significant potential in real-time multi-sensor perception tasks due to their event-driven and parameter-efficient characteristics.…