6 papers
Fine-grained Human Motion Understanding with Language Models
Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew +2
In this work, we propose \methodname, an LLM-based model for fine-grained human motion understanding that represents motion as a sequence of skeletal poses with explicit timestamps…
PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction
Zhi-Yi Lin, Thomas Markhorst, Jouh Yeong Chew +1
Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are limited to single-modality or…
MuPPet: Multi-person 2D-to-3D Pose Lifting
Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew +2
Multi-person social interactions are inherently built on coherence and relationships among all individuals within the group, making multi-person localization and body pose estimati…
GazeHTA: End-to-end Gaze Target Detection with Head-Target Association
Zhi-Yi Lin, Jouh Yeong Chew, Jan van Gemert +1
Precisely detecting which object a person is paying attention to is critical for human-robot interaction since it provides important cues for the next action from the human user. W…
Diffusion-Based Imitation Learning for Social Pose Generation
Antonio Lech Martin-Ozimek, Isuru Jayarathne, Su Larb Mon +1
Intelligent agents, such as robots and virtual agents, must understand the dynamics of complex social interactions to interact with humans. Effectively representing social dynamics…
Learning Nonverbal Cues in Multiparty Social Interactions for Robotic Facilitators
Antonio Lech Martin-Ozimek, Isuru Jayarathne, Su Larb Mon +1
Conventional behavior cloning (BC) models often struggle to replicate the subtleties of human actions. Previous studies have attempted to address this issue through the development…