collaborators

10 papers

cs.SD2026

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

Haotian Qi, Gabriel Skantze

Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduc…

cs.CL2026

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

Luca Modica, Filip Landin, Mehrdad Farahani +3

In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechan…

cs.CL2026

Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models

Yu Wang, Leyi Lao, Langchu Huang +3

Backchannels and fillers are important linguistic expressions in dialogue, but often treated as 'noise' to be bypassed in modern transformer-based language models (LMs). Here, we s…

eess.AS2026

VoXtream2: Full-stream TTS with dynamic speaking rate control

Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze

Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a…

cs.RO2025

Into the Wild: When Robots Are Not Welcome

Shaul Ashkenazi, Gabriel Skantze, Jane Stuart-Smith +1

Social robots are increasingly being deployed in public spaces, where they face not only technological difficulties and unexpected user utterances, but also objections from stakeho…

cs.CL2025

Detecting Referring Expressions in Visually Grounded Dialogue with Autoregressive Language Models

Bram Willemsen, Gabriel Skantze

In this paper, we explore the use of a text-only, autoregressive language modeling approach for the extraction of referring expressions from visually grounded dialogue. More specif…