8 papers · 1 filter
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by…
RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation
Sixu Lin, Junliang Chen, Huaiyuan Xu +8
Planning and acting in 3D environments is a fundamental capability for robotic manipulation in the real world. Although prior work has explored predictive flow planners to guide 3D…
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization
Sixu Lin, Yunpeng Qing, Litao Liu +4
Recent progress in Reinforcement Learning (RL) provides a principled approach to optimizing Vision-Language-Action (VLA) models, facilitating a shift from trajectory imitation to a…
SignBot: Learning Human-to-Humanoid Sign Language Interaction
Guanren Qiao, Sixu Lin, Ronglai Zuo +3
Sign language is a natural and visual form of language that uses movements and expressions to convey meaning, serving as a crucial means of communication for individuals who are de…
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
Shaopeng Zhai, Qi Zhang, Tianyi Zhang +7
Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLA…
SMAP: Self-supervised Motion Adaptation for Physically Plausible Humanoid Whole-body Control
Haoyu Zhao, Sixu Lin, Qingwei Ben +5
This paper presents a novel framework that enables real-world humanoid robots to maintain stability while performing human-like motion. Current methods train a policy which allows…