collaborators

5 papers

cs.RO2026

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Shijie Lian, Bin Yu, Xiaopeng Lin +8

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-hori…

cs.RO2026

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models

Bin Yu, Yao Zhang, Haishan Liu +9

Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction.…

cs.RO2026

PhysBrain 1.0 Technical Report

Shijie Lian, Bin Yu, Xiaopeng Lin +10

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a comple…

cs.RO2026

FrameSkip: Learning from Fewer but More Informative Frames in VLA Training

Bin Yu, Shijie Lian, Xiaopeng Lin +8

Vision-Language-Action (VLA) policies are commonly trained from dense robot demonstration trajectories, often collected through teleoperation, by sampling every recorded frame as i…

cs.RO2026

3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models

Bin Yu, Shijie Lian, Xiaopeng Lin +8

Vision-Language-Action (VLA) models leverage Multimodal Large Language Models (MLLMs) for robotic control, but recent studies reveal that MLLMs exhibit limited spatial intelligence…