collaborators
Showing cs.ROShow all

6 papers · 1 filter

cs.RO2026

SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation

Nie Lin, Takehiko Ohkawa, Sijin Chen +10

Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulatio…

cs.RO2026

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

Sijin Chen, Kaixuan Jiang, Haixin Shi +6

We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, abundant, and diverse, making it…

cs.RO2026

TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models

Wenbo Zhang, Jianxiong Li, Shuai Yang +4

Vision-Language-Action (VLA) models trained on large-scale data have made remarkable progress, but they remain vulnerable to distribution shifts at deployment time. Recent VLA mode…

cs.RO2026

From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors

Zhengshen Zhang, Hao Li, Yalun Dai +10

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptabilit…

cs.RO2026

World Guidance: World Modeling in Condition Space for Action Generation

Yue Su, Sijin Chen, Haixin Shi +7

Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, e…

cs.RO2025

GR-3 Technical Report

Chilam Cheang, Sijin Chen, Zhongren Cui +18

We report our recent progress towards building generalist robot policies, the development of GR-3. GR-3 is a large-scale vision-language-action (VLA) model. It showcases exceptiona…