activity
20162023
most citedGenerative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows

1 citations · 2 across the 3 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2023

VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation

Rohan Badlani, Akshit Arora, Subhankar Ghosh +5

We introduce VANI, a very lightweight multi-lingual accent controllable speech synthesis system. Our model builds upon disentanglement strategies proposed in RADMMM and supports ex…

cs.SD2023★ 1 cited

Multilingual Multiaccented Multispeaker TTS with RADTTS

Rohan Badlani, Rafael Valle, Kevin J. Shih +3

We work to create a multilingual speech synthesis system which can generate speech with the proper accent while retaining the characteristics of an individual voice. This is challe…

cs.SD2022★ 1 cited

Generative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows

Kevin J. Shih, Rafael Valle, Rohan Badlani +2

Despite recent advances in generative modeling for text-to-speech synthesis, these models do not yet have the same fine-grained adjustability of pitch-conditioned deterministic mod…

cs.SD2017

Speech Dereverberation with Context-aware Recurrent Neural Networks

Joao Felipe Santos, Tiago H. Falk

In this paper, we propose a model to perform speech dereverberation by estimating its spectral magnitude from the reverberant counterpart. Our models are capable of extracting feat…

cs.SD2017

Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency Mask

Stylianos Ioannis Mimilakis, Konstantinos Drossos, João F. Santos +3

Singing voice separation based on deep learning relies on the usage of time-frequency masking. In many cases the masking process is not a learnable function or is not encapsulated…

cs.SD2016

Music transcription modelling and composition using deep learning

Bob L. Sturm, João Felipe Santos, Oded Ben-Tal +1

We apply deep learning methods, specifically long short-term memory (LSTM) networks, to music transcription modelling and composition. We build and train LSTM networks using approx…