activity
20172026
most citedAnalyzing Input and Output Representations for Speech-Driven Gesture Generation

154 citations · 313 across the 18 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

VoXtream2: Full-stream TTS with dynamic speaking rate control

Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze

Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a…

eess.AS2025

When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs

Shree Harsha Bokkahalli Satish, Gustav Eje Henter, Éva Székely

The rapid development of SpeechLLM-based conversational AI systems has created a need for robustly benchmarking these efforts, including aspects of fairness and bias. At present, s…

eess.AS2025

VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency

Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze

We present VoXtream, a fully autoregressive, zero-shot streaming text-to-speech (TTS) system for real-time use that begins speaking from the first word. VoXtream directly maps inco…

eess.AS20224 cited

Wavebender GAN: An architecture for phonetically meaningful speech manipulation

Gustavo Teodoro Döhler Beck, Ulme Wennberg, Zofia Malisz +1

Deep learning has revolutionised synthetic speech quality. However, it has thus far delivered little value to the speech science community. The new methods do not meet the controll…

eess.AS2018

Analysing Shortcomings of Statistical Parametric Speech Synthesis

Gustav Eje Henter, Simon King, Thomas Merritt +1

Output from statistical parametric speech synthesis (SPSS) remains noticeably worse than natural speech recordings in terms of quality, naturalness, speaker similarity, and intelli…

eess.AS2018

Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis

Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang +1

Generating versatile and appropriate synthetic speech requires control over the output expression separate from the spoken text. Important non-textual speech variation is seldom an…