activity
20182020
most citedIn Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data

1 citations · 1 across the 2 of their papers we have counts for

collaborators

6 papers

eess.AS2020

Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech

Monica Sunkara, Srikanth Ronanki, Dhanush Bekal +2

In this work, we explore a multimodal semi-supervised learning approach for punctuation prediction by learning representations from large amounts of unlabelled audio and text data.…

cs.CL2020

Robust Prediction of Punctuation and Truecasing for Medical ASR

Monica Sunkara, Srikanth Ronanki, Kalpit Dixit +2

Automatic speech recognition (ASR) systems in the medical domain that focus on transcribing clinical dictations and doctor-patient conversations often pose many challenges due to t…

eess.AS2019

ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech

Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37

Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…

eess.AS2019

Fine-grained robust prosody transfer for single-speaker neural text-to-speech

Viacheslav Klimkov, Srikanth Ronanki, Jonas Rohnke +1

We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use eithe…

cs.CL20191 cited

In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data

Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote +5

Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes…

cs.CL2018

Effect of data reduction on sequence-to-sequence neural TTS

Javier Latorre, Jakub Lachowicz, Jaime Lorenzo-Trueba +4

Recent speech synthesis systems based on sampling from autoregressive neural networks models can generate speech almost undistinguishable from human recordings. However, these mode…