activity
20242026
collaborators

8 papers

cs.RO2026

LAMP: Latent Motion Prior-Guided Real-World Learning for Dexterous Hand Manipulation

Xinye Yang, Zhiyuan Ma, Hongze Yu +5

Real-world learning for dexterous hands remains brittle because high-dimensional hand actions amplify imitation errors and make reinforcement-learning exploration prone to contact-…

cs.CV2026

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

Hanqing Wang, Mingyu Liu, Xiaoyu Chen +9

3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordanc…

cs.RO2026

CADGrasp: Learning Contact and Collision Aware General Dexterous Grasping in Cluttered Scenes

Jiyao Zhang, Zhiyuan Ma, Tianhao Wu +2

Dexterous grasping in cluttered environments presents substantial challenges due to the high degrees of freedom of dexterous hands, occlusion, and potential collisions arising from…

cs.CV2025

Towards Cross-View Point Correspondence in Vision-Language Models

Yipu Wang, Yuheng Ji, Yuyang Liu +10

Cross-view correspondence is a fundamental capability for spatial understanding and embodied AI. However, it is still far from being realized in Vision-Language Models (VLMs), espe…

cs.CV2025

UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition

Guojun Lei, Rong Zhang, Chi Wang +4

We propose a novel architecture UniTransfer, which introduces both spatial and diffusion timestep decomposition in a progressive paradigm, achieving precise and controllable video…

cs.CV2025

Dream3DAvatar: Text-Controlled 3D Avatar Reconstruction from a Single Image

Gaofeng Liu, Hengsen Li, Ruoyu Gao +3

With the rapid advancement of 3D representation techniques and generative models, substantial progress has been made in reconstructing full-body 3D avatars from a single image. How…