8 papers · 1 filter
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
Shijie Lian, Bin Yu, Xiaopeng Lin +8
Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-hori…
SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models
Changti Wu, Bin Yu, Zhaolong Shen +6
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policie…
Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments
Xiaopeng Lin, Ruoqi Yang, Shijie Lian +14
Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot d…
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models
Bin Yu, Yao Zhang, Haishan Liu +9
Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction.…
PhysBrain 1.0 Technical Report
Shijie Lian, Bin Yu, Xiaopeng Lin +10
Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a comple…
FrameSkip: Learning from Fewer but More Informative Frames in VLA Training
Bin Yu, Shijie Lian, Xiaopeng Lin +8
Vision-Language-Action (VLA) policies are commonly trained from dense robot demonstration trajectories, often collected through teleoperation, by sampling every recorded frame as i…