132 citations · 157 across the 52 of their papers we have counts for
6 papers · 1 filter
Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos
Danze Chen, Yanzhe Chen, Qiming Huang +3
Vision-Language-Action (VLA) models require large-scale video-action pairs, yet real teleoperation remains scarce. While generated robot videos offer a scalable alternative, existi…
Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation
Kevin Yuchen Ma, Heng Zhang, Weisi Lin +2
Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Language-Action (VLA) models, often la…
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
Zechen Bai, Chen Gao, Mike Zheng Shou
Achieving truly adaptive embodied intelligence requires agents that learn not just by imitating static demonstrations, but by continuously improving through environmental interacti…
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
Hai Ci, Xiaokang Liu, Pei Yang +2
Robots that learn manipulation skills from everyday human videos could acquire broad capabilities without tedious robot data collection. We propose a video-to-video translation fra…
MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping
Heng Zhang, Kevin Yuchen Ma, Mike Zheng Shou +2
Dexterous grasping with multi-fingered hands remains challenging due to high-dimensional articulations and the cost of optimization-based pipelines. Existing end-to-end methods req…
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
Jianxin Bi, Kevin Yuchen Ma, Ce Hao +2
Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models lack the abi…