Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
arXiv:2509.12507 · doi:10.3389/frobt.2023.1110534
Abstract
One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on verbal modes such as language and speech, non-verbal communication is crucial for flexible interaction. We present a framework for generating pointing gestures in embodied agents by combining imitation and reinforcement learning. Using a small motion capture dataset, our method learns a motor control policy that produces physically valid, naturalistic gestures with high referential accuracy. We evaluate the approach against supervised learning and retrieval baselines in both objective metrics and a virtual reality referential game with human users. Results show that our system achieves higher naturalness and accuracy than state-of-the-art supervised models, highlighting the promise of imitation-RL for communicative gesture generation and its potential application to robots.
DOI: 10.3389/frobt.2023.1110534. This is the author's LaTeX version
References in corpus (14)
- NICE: Non-linear Independent Components Estimation
- Benchmarking Deep Reinforcement Learning for Continuous Control
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
- Emergence of Locomotion Behaviours in Rich Environments
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- MoGlow: Probabilistic and controllable motion synthesis using normalising flows
- Learning human behaviors from motion capture by adversarial imitation
- Transflower: probabilistic autoregressive dance generation with multimodal attention
- A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA Challenge 2020
- DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following
- Moving fast and slow: Analysis of representations and post-processing in speech-driven automatic gesture generation
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers
- Interactive Language: Talking to Robots in Real Time
- Imitation Learning of Robot Policies by Combining Language, Vision and Demonstration
Cited by in corpus (4)
- Diffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation
- Gesture Evaluation in Virtual Reality
- Exploring the Impact of Non-Verbal Virtual Agent Behavior on User Engagement in Argumentative Dialogues
- Incorporating Spatial Awareness in Data-Driven Gesture Generation for Virtual Agents