13 papers
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Yuehao Huang, Yunzi Wu, Xiaotao Zhang +7
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observati…
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
Zeyu Liu, Zhangzhe Zhu, Yang Zhang +3
Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-loop evaluation via simulator rollouts offers a more faithful assessme…
PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
Yang Zhang, Jiangyuan Zhao, Chenyou Fan +11
Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloni…
ReMoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement
Jiakun Zheng, Ting Xiao, Shiqin Cao +3
Text-to-motion (T2M) generation aims to control the behavior of a target character via textual descriptions. Leveraging text-motion paired datasets, existing T2M models have achiev…
Efficient Cross-Domain Offline Reinforcement Learning with Dynamics- and Value-Aligned Data Filtering
Zhongjian Qiao, Rui Yang, Jiafei Lyu +4
Cross-domain offline reinforcement learning (RL) aims to train a well-performing agent in the target environment, leveraging both a limited target domain dataset and a source domai…
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
Siyuan Yang, Yang Zhang, Haoran He +4
Vision-Language-Action (VLA) models, trained via flow-matching or diffusion objectives, excel at learning complex behaviors from large-scale, multi-modal datasets (e.g., human tele…