activity
20242026
collaborators

7 papers

cs.CV2026

Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?

Masatoshi Tateno, Alexandros Stergiou, Risa Shinoda +2

Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts b…

cs.AI2025

EgoBrain: Synergizing Minds and Eyes For Human Action Understanding

Nie Lin, Yansen Wang, Dongqi Han +5

The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human co…

cs.CV2025

EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking

Yuki Sakai, Ryosuke Furuta, Juichun Yen +1

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill trans…

cs.CV2025

Leadership Assessment in Pediatric Intensive Care Unit Team Training

Liangyang Ouyang, Yuki Sakai, Ryosuke Furuta +3

This paper addresses the task of assessing PICU team's leadership skills by developing an automated analysis framework based on egocentric vision. We identify key behavioral cues,…

cs.CV2024

Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos

Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura +4

We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos with exocentric views to an egocentric view.…

cs.CV2024

Learning Multiple Object States from Actions via Large Language Models

Masatoshi Tateno, Takuma Yagi, Ryosuke Furuta +1

Recognizing the states of objects in a video is crucial in understanding the scene beyond actions and objects. For instance, an egg can be raw, cracked, and whisked while cooking a…