activity
20232025
most citedDiffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation

23 citations · 37 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2024

Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech

Shivam Mehta, Harm Lameris, Rajiv Punmiya +3

Converting input symbols to output audio in TTS requires modelling the durations of speech sounds. Leading non-autoregressive (NAR) TTS models treat duration modelling as a regress…

eess.AS2023

Unified speech and gesture synthesis using flow matching

Shivam Mehta, Ruibo Tu, Simon Alexanderson +3

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviou…

eess.AS2023★ 23 cited

Diffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation

Anna Deichler, Shivam Mehta, Simon Alexanderson +1

This paper describes a system developed for the GENEA (Generation and Evaluation of Non-verbal Behaviour for Embodied Agents) Challenge 2023. Our solution builds on an existing dif…

eess.AS2023★ 1 cited

Matcha-TTS: A fast TTS architecture with conditional flow matching

Shivam Mehta, Ruibo Tu, Jonas Beskow +2

We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an OD…

eess.AS2023★ 13 cited

Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis

Shivam Mehta, Siyang Wang, Simon Alexanderson +3

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face…