1 citations · 1 across the 2 of their papers we have counts for
6 papers
Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech
Monica Sunkara, Srikanth Ronanki, Dhanush Bekal +2
In this work, we explore a multimodal semi-supervised learning approach for punctuation prediction by learning representations from large amounts of unlabelled audio and text data.…
Robust Prediction of Punctuation and Truecasing for Medical ASR
Monica Sunkara, Srikanth Ronanki, Kalpit Dixit +2
Automatic speech recognition (ASR) systems in the medical domain that focus on transcribing clinical dictations and doctor-patient conversations often pose many challenges due to t…
ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37
Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…
Fine-grained robust prosody transfer for single-speaker neural text-to-speech
Viacheslav Klimkov, Srikanth Ronanki, Jonas Rohnke +1
We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use eithe…
In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data
Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote +5
Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes…
Effect of data reduction on sequence-to-sequence neural TTS
Javier Latorre, Jakub Lachowicz, Jaime Lorenzo-Trueba +4
Recent speech synthesis systems based on sampling from autoregressive neural networks models can generate speech almost undistinguishable from human recordings. However, these mode…