8 papers
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
Anna Deichler
A robot resolving ``put the cup on that one'' must fuse gesture, language, and scene geometry, yet 3D grounding benchmarks only partially capture this regime: descriptions are writ…
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan +4
Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel…
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
Anna Deichler, Jonas Beskow
We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and sta…
Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation
Anna Deichler, Siyang Wang, Simon Alexanderson +1
Pointing is a key mode of interaction with robots, yet most prior work has focused on recognition rather than generation. We present a motion capture dataset of human pointing gest…
Gesture Evaluation in Virtual Reality
Axel Wiebe Werner, Jonas Beskow, Anna Deichler
Gestures are central to human communication, enriching interactions through non-verbal expression. Virtual avatars increasingly use AI-generated gestures to enhance life-likeness,…
Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
Anna Deichler, Siyang Wang, Simon Alexanderson +1
One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on ve…