5 papers
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan +4
Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel…
Multi-Axis Speech Similarity via Factor-Partitioned Embeddings
Jim O'Regan, Jens Edlund
Speech encodes multiple simultaneous attributes -- linguistic content, speaker identity, dialect, gender --that conventional single-vector embeddings conflate. We present a factor-…
Grounded Gesture Generation: Language, Motion, and Space
Anna Deichler, Jim O'Regan, Teo Guichoux +2
Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing…
MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans
Anna Deichler, Jim O'Regan, Jonas Beskow
In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is t…
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
Shivam Mehta, Anna Deichler, Jim O'Regan +4
Although humans engaged in face-to-face conversation simultaneously communicate both verbally and non-verbally, methods for joint and unified synthesis of speech audio and co-speec…