4 papers · 1 filter
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan +4
Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel…
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
Anna Deichler, Jonas Beskow
We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and sta…
Grounded Gesture Generation: Language, Motion, and Space
Anna Deichler, Jim O'Regan, Teo Guichoux +2
Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing…
MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans
Anna Deichler, Jim O'Regan, Jonas Beskow
In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is t…