8 papers
Robot Learning from Human Videos: A Survey
Junyi Ma, Erhang Zhang, Haoran Yang +4
A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation sk…
EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
Junyi Ma, Erhang Zhang, Yin-Dong Zheng +3
Analyzing hand-object interaction in egocentric vision facilitates VR/AR applications and human-robot policy transfer. Existing research has mostly focused on modeling the behavior…
Zero-Shot Temporal Interaction Localization for Egocentric Videos
Erhang Zhang, Junyi Ma, Yin-Dong Zheng +2
Locating human-object interaction (HOI) actions within video serves as the foundation for multiple downstream tasks, such as human behavior analysis and human-robot skill transfer.…
Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction
Junyi Ma, Wentao Bao, Jingyi Xu +3
Predicting hand motion is critical for understanding human intentions and bridging the action space between human movements and robot manipulations. Existing hand trajectory predic…
EADReg: Probabilistic Correspondence Generation with Efficient Autoregressive Diffusion Model for Outdoor Point Cloud Registration
Linrui Gong, Jiuming Liu, Junyi Ma +3
Diffusion models have shown the great potential in the point cloud registration (PCR) task, especially for enhancing the robustness to challenging cases. However, existing diffusio…
Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting
Jingyi Xu, Xieyuanli Chen, Junyi Ma +4
The task of occupancy forecasting (OCF) involves utilizing past and present perception data to predict future occupancy states of autonomous vehicle surrounding environments, which…