collaborators

18 papers

cs.RO2026

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Shijie Lian, Bin Yu, Xiaopeng Lin +8

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-hori…

cs.RO2026

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

Changti Wu, Bin Yu, Zhaolong Shen +6

Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policie…

cs.RO2026

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments

Xiaopeng Lin, Ruoqi Yang, Shijie Lian +14

Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot d…

cs.CV2026

TUGS: Physics-based Compact Representation of Underwater Scenes by Tensorized Gaussian

Shijie Lian, Ziyi Zhang, Hua Li +4

Underwater 3D scene reconstruction is crucial for multimedia applications in adverse environments, such as underwater robotic perception and navigation. However, the complexity of…

cs.RO2026

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models

Bin Yu, Yao Zhang, Haishan Liu +9

Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction.…

cs.AI2026

LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

Shijie Lian, Bin Yu, Xiaopeng Lin +6

Vision-Language-Action (VLA) models have shown promise in robot manipulation but often struggle to generalize to new instructions or complex multi-task scenarios. We identify a cri…