5 papers
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
Runze Li, Hongyin Zhang, Junxi Jin +5
Vision-Language-Action (VLA) models have emerged as a promising paradigm for building embodied agents that ground perception and language into action. However, most existing approa…
CMR: Contractive Mapping Embeddings for Robust Humanoid Locomotion on Unstructured Terrains
Qixin Zeng, Hongyin Zhang, Shangke Lyu +3
Robust disturbance rejection remains a longstanding challenge in humanoid locomotion, particularly on unstructured terrains where sensing is unreliable and model mismatch is pronou…
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
Hongyin Zhang, Shuo Zhang, Junxi Jin +3
Vision-Language-Action (VLA) models have recently emerged as powerful general-purpose policies for robotic manipulation, benefiting from large-scale multi-modal pre-training. Howev…
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
Hongyin Zhang, Shiyuan Zhang, Junxi Jin +4
Vision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these mod…
PEP-GS: Perceptually-Enhanced Precise Structured 3D Gaussians for View-Adaptive Rendering
Junxi Jin, Xiulai Li, Haiping Huang +3
Recently, 3D Gaussian Splatting (3D-GS) has achieved significant success in real-time, high-quality 3D scene rendering. However, it faces several challenges, including Gaussian red…