From the 1 of 14 linked papers with an AI index.
13 papers · 1 filter
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
Naru Suzuki, Takehiko Ohkawa, Tatsuro Banno +3
The paper presents a diffusion-based generative prior that refines 3D hand pose reconstruction by using affordance-aware textual descriptions of hand‑object interactions, improving…
Multi-speaker Attention Alignment for Multimodal Social Interaction
Liangyang Ouyang, Yifei Huang, Mingfang Zhang +3
Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While…
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
Tatsuro Banno, Takehiko Ohkawa, Ruicong Liu +2
Bimanual human activities inherently involve coordinated movements of both hands and body. However, the impact of this coordination in activity understanding has not been systemati…
Generative Modeling of Shape-Dependent Self-Contact Human Poses
Takehiko Ohkawa, Jihyun Lee, Shunsuke Saito +7
One can hardly model self-contact of human poses without considering underlying body shapes. For example, the pose of rubbing a belly for a person with a low BMI leads to penetrati…
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
Yuki Sakai, Ryosuke Furuta, Juichun Yen +1
Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill trans…
Leadership Assessment in Pediatric Intensive Care Unit Team Training
Liangyang Ouyang, Yuki Sakai, Ryosuke Furuta +3
This paper addresses the task of assessing PICU team's leadership skills by developing an automated analysis framework based on egocentric vision. We identify key behavioral cues,…