6 citations · 9 across the 6 of their papers we have counts for
10 papers
Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
Livia Qian, Gabriel Skantze
Short feedback responses, such as backchannels, play an important role in spoken dialogue. So far, most of the modeling of feedback responses has focused on their timing, often neg…
Multilingual Turn-taking Prediction Using Voice Activity Projection
Koji Inoue, Bing'er Jiang, Erik Ekstedt +2
This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin…
An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems
Koji Inoue, Divesh Lala, Keiko Ochi +2
Establishing evaluation schemes for spoken dialogue systems is important, but it can also be challenging. While subjective evaluations are commonly used in user experiments, object…
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
Koji Inoue, Bing'er Jiang, Erik Ekstedt +2
A demonstration of a real-time and continuous turn-taking prediction system is presented. The system is based on a voice activity projection (VAP) model, which directly maps dialog…
Towards Objective Evaluation of Socially-Situated Conversational Robots: Assessing Human-Likeness through Multimodal User Behaviors
Koji Inoue, Divesh Lala, Keiko Ochi +2
This paper tackles the challenging task of evaluating socially situated conversational robots and presents a novel objective evaluation approach that relies on multimodal user beha…
Resolving References in Visually-Grounded Dialogue via Text Generation
Bram Willemsen, Livia Qian, Gabriel Skantze
Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge…