most citedUsing Large Language Models for Zero-Shot Natural Language Generation from Knowledge Graphs

6 citations · 9 across the 6 of their papers we have counts for

collaborators

10 papers

cs.CL20241 cited

Joint Learning of Context and Feedback Embeddings in Spoken Dialogue

Livia Qian, Gabriel Skantze

Short feedback responses, such as backchannels, play an important role in spoken dialogue. So far, most of the modeling of feedback responses has focused on their timing, often neg…

cs.CL20241 cited

Multilingual Turn-taking Prediction Using Voice Activity Projection

Koji Inoue, Bing'er Jiang, Erik Ekstedt +2

This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin…

cs.CL2024

An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems

Koji Inoue, Divesh Lala, Keiko Ochi +2

Establishing evaluation schemes for spoken dialogue systems is important, but it can also be challenging. While subjective evaluations are commonly used in user experiments, object…

cs.CL20242 cited

Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection

Koji Inoue, Bing'er Jiang, Erik Ekstedt +2

A demonstration of a real-time and continuous turn-taking prediction system is presented. The system is based on a voice activity projection (VAP) model, which directly maps dialog…

cs.CL2023

Towards Objective Evaluation of Socially-Situated Conversational Robots: Assessing Human-Likeness through Multimodal User Behaviors

Koji Inoue, Divesh Lala, Keiko Ochi +2

This paper tackles the challenging task of evaluating socially situated conversational robots and presents a novel objective evaluation approach that relies on multimodal user beha…

cs.CL2023

Resolving References in Visually-Grounded Dialogue via Text Generation

Bram Willemsen, Livia Qian, Gabriel Skantze

Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge…