From the 1 of 8 linked papers with an AI index.
8 papers
EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation
Xinyuan Guan, Feifan Chen, Xinyu Zhan +3
Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability to complex tasks calls for c…
Multi-view Hand Reconstruction with a Point-Embedded Transformer
Lixin Yang, Licheng Zhong, Pengxiang Zhu +4
The paper presents POEM, a multi-view hand mesh reconstruction system that embeds static basis points in the multi-view stereo space and uses a transformer to fuse features across…
ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning
Bokai Lin, Yifu Xu, Xinyu Zhan +6
Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past expe…
TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning
Jiacheng Liu, Pengxiang Ding, Qihang Zhou +8
Recent Vision-Language-Action models show potential to generalize across embodiments but struggle to quickly align with a new robot's action space when high-quality demonstrations…
VFM-Recon: Unlocking Cross-Domain Scene-Level Neural Reconstruction with Scale-Aligned Foundation Priors
Yuhang Ming, Tingkang Xi, Xingrui Yang +4
Scene-level neural volumetric reconstruction from monocular videos remains challenging, especially under severe domain shifts. Although recent advances in vision foundation models…
Motion Before Action: Diffusing Object Motion as Manipulation Condition
Yue Su, Xinyu Zhan, Hongjie Fang +3
Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning…