1 citations · 2 across the 3 of their papers we have counts for
6 papers · 1 filter
VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation
Rohan Badlani, Akshit Arora, Subhankar Ghosh +5
We introduce VANI, a very lightweight multi-lingual accent controllable speech synthesis system. Our model builds upon disentanglement strategies proposed in RADMMM and supports ex…
Multilingual Multiaccented Multispeaker TTS with RADTTS
Rohan Badlani, Rafael Valle, Kevin J. Shih +3
We work to create a multilingual speech synthesis system which can generate speech with the proper accent while retaining the characteristics of an individual voice. This is challe…
Generative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows
Kevin J. Shih, Rafael Valle, Rohan Badlani +2
Despite recent advances in generative modeling for text-to-speech synthesis, these models do not yet have the same fine-grained adjustability of pitch-conditioned deterministic mod…
Speech Dereverberation with Context-aware Recurrent Neural Networks
Joao Felipe Santos, Tiago H. Falk
In this paper, we propose a model to perform speech dereverberation by estimating its spectral magnitude from the reverberant counterpart. Our models are capable of extracting feat…
Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency Mask
Stylianos Ioannis Mimilakis, Konstantinos Drossos, João F. Santos +3
Singing voice separation based on deep learning relies on the usage of time-frequency masking. In many cases the masking process is not a learnable function or is not encapsulated…
Music transcription modelling and composition using deep learning
Bob L. Sturm, João Felipe Santos, Oded Ben-Tal +1
We apply deep learning methods, specifically long short-term memory (LSTM) networks, to music transcription modelling and composition. We build and train LSTM networks using approx…