Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
Bohan Zhou, Yi Zhan, Zhongbin Zhang +1
Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, self-occlusions, perspective disto…
cs.CV2024
NOLO: Navigate Only Look Once
Bohan Zhou, Zhongbin Zhang, Jiangxing Wang +1
The in-context learning ability of Transformer models has brought new possibilities to visual navigation. In this paper, we focus on the video navigation setting, where an in-conte…
cs.CV2024
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
Hao Luo, Bohan Zhou, Zongqing Lu
Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amoun…