8 papers
Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study
Koji Inoue, Mikey Elmers, Yahui Fu +5
We present a multilingual, continuous backchannel prediction model for Japanese, English, and Chinese, and use it to investigate cross-linguistic timing behavior. The model is Tran…
Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems
Mikey Elmers, Koji Inoue, Divesh Lala +1
Turn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP…
Prompt-Guided Turn-Taking Prediction
Koji Inoue, Mikey Elmers, Yahui Fu +4
Turn-taking prediction models are essential components in spoken dialogue systems and conversational robots. Recent approaches leverage transformer-based architectures to predict s…
An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
Koji Inoue, Divesh Lala, Mikey Elmers +2
Handling multi-party dialogues represents a significant step for advancing spoken dialogue systems, necessitating the development of tasks specific to multi-party interactions. To…
Why Do We Laugh? Annotation and Taxonomy Generation for Laughable Contexts in Spontaneous Text Conversation
Koji Inoue, Mikey Elmers, Divesh Lala +1
Laughter serves as a multifaceted communicative signal in human interaction, yet its identification within dialogue presents a significant challenge for conversational AI systems.…
Does the Appearance of Autonomous Conversational Robots Affect User Spoken Behaviors in Real-World Conference Interactions?
Zi Haur Pang, Yahui Fu, Divesh Lala +3
We investigate the impact of robot appearance on users' spoken behavior during real-world interactions by comparing a human-like android, ERICA, with a less anthropomorphic humanoi…