activity
20172022
most citedLearning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks

2 citations · 2 across the 3 of their papers we have counts for

collaborators

9 papers

eess.AS2022

Distribution augmentation for low-resource expressive text-to-speech

Mateusz Lajszczak, Animesh Prasad, Arent van Korlaar +8

This paper presents a novel data augmentation technique for text-to-speech (TTS), that allows to generate new (text, audio) training examples without requiring any additional data.…

cs.CL2021

Proteno: Text Normalization with Limited Data for Fast Deployment in Text to Speech Systems

Shubhi Tyagi, Antonio Bonafonte, Jaime Lorenzo-Trueba +1

Developing Text Normalization (TN) systems for Text-to-Speech (TTS) on new languages is hard. We propose a novel architecture to facilitate it for multiple languages while using da…

cs.CL2019

Prosodic Phrase Alignment for Machine Dubbing

Alp Öktem, Mireia Farrús, Antonio Bonafonte

Dubbing is a type of audiovisual translation where dialogues are translated and enacted so that they give the impression that the media is in the target language. It requires a car…

cs.SD2019

Problem-Agnostic Speech Embeddings for Multi-Speaker Text-to-Speech with SampleRNN

David Álvarez, Santiago Pascual, Antonio Bonafonte

Text-to-speech (TTS) acoustic models map linguistic features into an acoustic representation out of which an audible waveform is generated. The latest and most natural TTS systems…

cs.SD2019

Towards Generalized Speech Enhancement with Generative Adversarial Networks

Santiago Pascual, Joan Serrà, Antonio Bonafonte

The speech enhancement task usually consists of removing additive noise or reverberation that partially mask spoken utterances, affecting their intelligibility. However, little att…

cs.LG20192 cited

Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks

Santiago Pascual, Mirco Ravanelli, Joan Serrà +2

Learning good representations without supervision is still an open issue in machine learning, and is particularly challenging for speech signals, which are often characterized by l…