514 citations · 643 across the 6 of their papers we have counts for
9 papers
KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE using Mel-spectrograms
Chien-Feng Liao, Jen-Yu Liu, Yi-Hsuan Yang
In this paper, we propose a novel neural network model called KaraSinger for a less-studied singing voice synthesis (SVS) task named score-free SVS, in which the prosody and melody…
SpeechBrain: A General-Purpose Speech Toolkit
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga +18
SpeechBrain is an open-source and all-in-one speech toolkit. It is designed to facilitate the research and development of neural speech processing technologies by being simple, fle…
Transformers with Competitive Ensembles of Independent Mechanisms
Alex Lamb, Di He, Anirudh Goyal +4
An important development in deep learning from the earliest MLPs has been a move towards architectures with structural inductive biases which enable the model to keep distinct sour…
Incorporating Broad Phonetic Information for Speech Enhancement
Yen-Ju Lu, Chien-Feng Liao, Xugang Lu +2
In noisy conditions, knowing speech contents facilitates listeners to more effectively suppress background noise components and to retrieve pure speech signals. Previous studies ha…
Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing
Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh +9
The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our st…
MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement
Szu-Wei Fu, Chien-Feng Liao, Yu Tsao +1
Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the…