6 papers
\textsc{Gen2Real}: Towards Demo-Free Dexterous Manipulation by Harnessing Generated Video
Kai Ye, Yuhang Wu, Shuyuan Hu +4
Dexterous manipulation remains a challenging robotics problem, largely due to the difficulty of collecting extensive human demonstrations for learning. In this paper, we introduce…
Grasp What You Want: Embodied Dexterous Grasping System Driven by Your Voice
Junliang Li, Kai Ye, Haolan Kang +6
In recent years, as robotics has advanced, human-robot collaboration has gained increasing importance. However, current robots struggle to fully and accurately interpret human inte…
A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language Model
Panwen Hu, Nan Xiao, Feifei Li +2
In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for…
Divide and Conquer: Improving Multi-Camera 3D Perception with 2D Semantic-Depth Priors and Input-Dependent Queries
Qi Song, Qingyong Hu, Chi Zhang +2
3D perception tasks, such as 3D object detection and Bird's-Eye-View (BEV) segmentation using multi-camera images, have drawn significant attention recently. Despite the fact that…
ClickAttention: Click Region Similarity Guided Interactive Segmentation
Long Xu, Shanghong Li, Yongquan Chen +3
Interactive segmentation algorithms based on click points have garnered significant attention from researchers in recent years. However, existing studies typically use sparse click…
MST: Adaptive Multi-Scale Tokens Guided Interactive Segmentation
Long Xu, Shanghong Li, Yongquan Chen +2
Interactive segmentation has gained significant attention for its application in human-computer interaction and data annotation. To address the target scale variation issue in inte…