82 citations · 254 across the 42 of their papers we have counts for
5 papers · 1 filter
Dynamic Prosody Generation for Speech Synthesis using Linguistics-Driven Acoustic Embedding Selection
Shubhi Tyagi, Marco Nicolis, Jonas Rohnke +2
Recent advances in Text-to-Speech (TTS) have improved quality and naturalness to near-human capabilities when considering isolated sentences. But something which is still lacking i…
In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data
Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote +5
Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes…
Active and Semi-Supervised Learning in ASR: Benefits on the Acoustic and Language Models
Thomas Drugman, Janne Pylkkonen, Reinhard Kneser
The goal of this paper is to simulate the benefits of jointly applying active learning (AL) and semi-supervised training (SST) in a new speech recognition application. Our data sel…
Effect of data reduction on sequence-to-sequence neural TTS
Javier Latorre, Jakub Lachowicz, Jaime Lorenzo-Trueba +4
Recent speech synthesis systems based on sampling from autoregressive neural networks models can generate speech almost undistinguishable from human recordings. However, these mode…
LSTM-based Whisper Detection
Zeynab Raeesy, Kellen Gillespie, Zhenpei Yang +6
This article presents a whisper speech detector in the far-field domain. The proposed system consists of a long-short term memory (LSTM) neural network trained on log-filterbank en…