From the 1 of 18 linked papers with an AI index.
18 papers
EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation
Xinyuan Guan, Feifan Chen, Xinyu Zhan +3
Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability to complex tasks calls for c…
Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies
Chenyi Wang, Xinkai Wang, Bokai Lin +4
Action labels tell a vision-language-action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains th…
Multi-view Hand Reconstruction with a Point-Embedded Transformer
Lixin Yang, Licheng Zhong, Pengxiang Zhu +4
The paper presents POEM, a multi-view hand mesh reconstruction system that embeds static basis points in the multi-view stereo space and uses a transformer to fuse features across…
AnyDexRT: Calibration-Free Dexterous Hand Retargeting with Few-Shot Human Guidance
Chenxi Wang, Ying Feng, Hongjie Fang +4
Teleoperation is a key interface for controlling dexterous robotic hands and collecting demonstrations for imitation learning. Its effectiveness largely depends on kinematic retarg…
ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning
Bokai Lin, Yifu Xu, Xinyu Zhan +6
Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past expe…
LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior
Xinkai Wang, Chenyi Wang, Yifu Xu +7
We introduce \textbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipulation.Existing VLA models reg…