6 papers
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Téo Guichoux, Théodor Lemerle, Shivam Mehta +5
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakeni…
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
Anna Deichler, Jonas Beskow
We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and sta…
Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation
Anna Deichler, Siyang Wang, Simon Alexanderson +1
Pointing is a key mode of interaction with robots, yet most prior work has focused on recognition rather than generation. We present a motion capture dataset of human pointing gest…
Gesture Evaluation in Virtual Reality
Axel Wiebe Werner, Jonas Beskow, Anna Deichler
Gestures are central to human communication, enriching interactions through non-verbal expression. Virtual avatars increasingly use AI-generated gestures to enhance life-likeness,…
Grounded Gesture Generation: Language, Motion, and Space
Anna Deichler, Jim O'Regan, Teo Guichoux +2
Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing…
MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans
Anna Deichler, Jim O'Regan, Jonas Beskow
In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is t…