activity
20212026
most citedIs Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

132 citations · 157 across the 52 of their papers we have counts for

collaborators
Showing cs.ROShow all

6 papers · 1 filter

cs.RO2026

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos

Danze Chen, Yanzhe Chen, Qiming Huang +3

Vision-Language-Action (VLA) models require large-scale video-action pairs, yet real teleoperation remains scarce. While generated robot videos offer a scalable alternative, existi…

cs.RO2026

Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation

Kevin Yuchen Ma, Heng Zhang, Weisi Lin +2

Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Language-Action (VLA) models, often la…

cs.RO2025

EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models

Zechen Bai, Chen Gao, Mike Zheng Shou

Achieving truly adaptive embodied intelligence requires agents that learn not just by imitating static demonstrations, but by continuously improving through environmental interacti…

cs.RO2025

H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos

Hai Ci, Xiaokang Liu, Pei Yang +2

Robots that learn manipulation skills from everyday human videos could acquire broad capabilities without tedious robot data collection. We propose a video-to-video translation fra…

cs.RO2025

MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping

Heng Zhang, Kevin Yuchen Ma, Mike Zheng Shou +2

Dexterous grasping with multi-fingered hands remains challenging due to high-dimensional articulations and the cost of optimization-based pipelines. Existing end-to-end methods req…

cs.RO2025

VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback

Jianxin Bi, Kevin Yuchen Ma, Ce Hao +2

Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models lack the abi…