21 papers
ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning
Wei Xiao, Weiliang Tang, Yuying Ge +4
Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models. However, enabling seamless humanoid interventions is a formidable syst…
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models
Lu Qiu, Yizhuo Li, Yi Chen +3
World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However…
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
Yi Chen, Yuying Ge, Hui Zhou +3
The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat…
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling
Boyu Chen, Yi Chen, Lu Qiu +3
Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment…
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
Jun Zhang, Teng Wang, Yuying Ge +4
This paper does not introduce a novel method but instead establishes a straightforward, incremental, yet essential baseline for video temporal grounding (VTG), a core capability in…
Aligning Latent Spaces with Flow Priors
Yizhuo Li, Yuying Ge, Yixiao Ge +2
This paper presents a novel framework for aligning learnable latent spaces to arbitrary target distributions by leveraging flow-based generative models as priors. Our method first…