PEAR: Phrase-Based Hand-Object Interaction Anticipation
arXiv:2407.21510 · doi:10.1007/s11432-024-4405-4
Abstract
First-person hand-object interaction anticipation aims to predict the interaction process over a forthcoming period based on current scenes and prompts. This capability is crucial for embodied intelligence and human-robot collaboration. The complete interaction process involves both pre-contact interaction intention (i.e., hand motion trends and interaction hotspots) and post-contact interaction manipulation (i.e., manipulation trajectories and hand poses with contact). Existing research typically anticipates only interaction intention while neglecting manipulation, resulting in incomplete predictions and an increased likelihood of intention errors due to the lack of manipulation constraints. To address this, we propose a novel model, PEAR (Phrase-Based Hand-Object Interaction Anticipation), which jointly anticipates interaction intention and manipulation. To handle uncertainties in the interaction process, we employ a twofold approach. Firstly, we perform cross-alignment of verbs, nouns, and images to reduce the diversity of hand movement patterns and object functional attributes, thereby mitigating intention uncertainty. Secondly, we establish bidirectional constraints between intention and manipulation using dynamic integration and residual connections, ensuring consistency among elements and thus overcoming manipulation uncertainty. To rigorously evaluate the performance of the proposed model, we collect a new task-relevant dataset, EGO-HOIP, with comprehensive annotations. Extensive experimental results demonstrate the superiority of our method.
22 pages, 10 figures, 4 tables
References in corpus (23)
- Learning Transferable Visual Models From Natural Language Supervision
- Embodied Hands: Modeling and Capturing Hands and Bodies Together
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Rolling-Unrolling LSTMs for Action Anticipation from First-Person Video
- Learning to Anticipate Egocentric Actions by Imagination
- LISA: Reasoning Segmentation via Large Language Model
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- Interacting Attention Graph for Single Image Two-Hand Reconstruction
- AlignSDF: Pose-Aligned Signed Distance Fields for Hand-Object Reconstruction
- MobRecon: Mobile-Friendly Hand Mesh Reconstruction from Monocular Image
- Hand-Object Contact Consistency Reasoning for Human Grasps Generation
- 3D Interacting Hand Pose Estimation by Hand De-occlusion and Removal
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos
- AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose Estimation
- Dexterous Functional Pre-Grasp Manipulation with Diffusion Policy
- Bidirectional Progressive Transformer for Interaction Intention Anticipation
- RBP-Pose: Residual Bounding Box Projection for Category-Level Pose Estimation
- Diff-IP2D: Diffusion-Based Hand-Object Interaction Prediction on Egocentric Videos
- OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion
- Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
- Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting
- LEMON: Learning 3D Human-Object Interaction Relation from 2D Images