activity
20202026
most citedBiHand: Recovering Hand Mesh with Multi-stage Bisected Hourglass Networks

26 citations · 67 across the 24 of their papers we have counts for

collaborators
Showing cs.ROShow all

11 papers · 1 filter

cs.RO2026

Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies

Chenyi Wang, Xinkai Wang, Bokai Lin +4

Action labels tell a vision-language-action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains th…

cs.RO2026

ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

Bokai Lin, Yifu Xu, Xinyu Zhan +6

Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past expe…

cs.RO2026

AnyDexRT: Calibration-Free Dexterous Hand Retargeting with Few-Shot Human Guidance

Chenxi Wang, Ying Feng, Hongjie Fang +4

Teleoperation is a key interface for controlling dexterous robotic hands and collecting demonstrations for imitation learning. Its effectiveness largely depends on kinematic retarg…

cs.RO2026

Revisiting Articulated Parts Perception in Robot Manipulation

Xiaoqian Wu, Yejie Guo, Xiaoyang Chen +3

We are surrounded by various objects with movable, articulated parts, e.g., box, handle, door. An accurate and generalizable perception of articulated parts is essential to enhance…

cs.RO2026

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction

Kai Xiong, Hongjie Fang, Lixin Yang +1

Effectively handling the interplay between spatial perception and action generation remains a critical bottleneck in robotic manipulation. Existing methods typically treat spatial…

cs.RO2026

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment

Yifu Xu, Bokai Lin, Xinyu Zhan +4

Scaling up robot learning is hindered by the scarcity of robotic demonstrations, whereas human videos offer a vast, untapped source of interaction data. However, bridging the embod…