activity
20242026
collaborators
Showing cs.ROShow all

6 papers · 1 filter

cs.RO2026

GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

Yanping Zhao, Hang Yu, Yiwei Wang +7

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promisin…

cs.RO2026

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

Shengzhuo Yang, Ronghao Yu, Chuanjie Lv +5

Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. R…

cs.RO2026

Learning Native Continuation for Action Chunking Flow Policies

Yufeng Liu, Hang Yu, Juntu Zhao +9

Action chunking enables Vision Language Action (VLA) models to run in real time, but naive chunked execution often exhibits discontinuities at chunk boundaries. Real-Time Chunking…

cs.RO2026

Robust and Generalized Humanoid Motion Tracking

Yubiao Ma, Han Yu, Jiayin Xie +7

Learning a general humanoid whole-body controller is challenging because practical reference motions can exhibit noise and inconsistencies after being transferred to the robot doma…

cs.RO2026

Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation

Di Zhang, Weicheng Duan, Dasen Gu +5

Real-world robotic manipulation demands visuomotor policies capable of robust spatial scene understanding and strong generalization across diverse camera viewpoints. While recent a…

cs.RO2025

DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot Handover

Youzhuo Wang, Jiayi Ye, Chuyang Xiao +6

Handover between a human and a dexterous robotic hand is a fundamental yet challenging task in human-robot collaboration. It requires handling dynamic environments and a wide varie…