8 papers
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
Ruiteng Zhao, Zhengshen Zhang, Yue Su +6
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both ali…
CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipation
Zhishan Tao, Ruoyu Wang, Yucheng Wu +6
Collision anticipation in autonomous driving requires not only accurate early warnings but also interpretable reasoning about what risk factors are being tracked and how risk evolv…
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
Chubin Zhang, Jianan Wang, Zifeng Gao +5
Generalist Vision-Language-Action models remain constrained by the scarcity of robotic data relative to the abundance of human video demonstrations. Existing Latent Action Models a…
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse
Kuan Zhang, Dongchen Liu, Qiyue Zhao +12
The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this singular physical existence…
World Guidance: World Modeling in Condition Space for Action Generation
Yue Su, Sijin Chen, Haixin Shi +7
Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, e…
DSPv2: Improved Dense Policy for Effective and Generalizable Whole-body Mobile Manipulation
Yue Su, Chubin Zhang, Sijin Chen +4
Learning whole-body mobile manipulation via imitation is essential for generalizing robotic skills to diverse environments and complex tasks. However, this goal is hindered by sign…