10 papers
MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild
Haotian Qi, Gabriel Skantze
Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduc…
Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?
Luca Modica, Filip Landin, Mehrdad Farahani +3
In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechan…
Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models
Yu Wang, Leyi Lao, Langchu Huang +3
Backchannels and fillers are important linguistic expressions in dialogue, but often treated as 'noise' to be bypassed in modern transformer-based language models (LMs). Here, we s…
VoXtream2: Full-stream TTS with dynamic speaking rate control
Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze
Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a…
Into the Wild: When Robots Are Not Welcome
Shaul Ashkenazi, Gabriel Skantze, Jane Stuart-Smith +1
Social robots are increasingly being deployed in public spaces, where they face not only technological difficulties and unexpected user utterances, but also objections from stakeho…
Detecting Referring Expressions in Visually Grounded Dialogue with Autoregressive Language Models
Bram Willemsen, Gabriel Skantze
In this paper, we explore the use of a text-only, autoregressive language modeling approach for the extraction of referring expressions from visually grounded dialogue. More specif…