6 citations · 6 across the 2 of their papers we have counts for
4 papers
Silence is Sweeter Than Speech: Self-Supervised Model Using Silence to Store Speaker Information
Chi-Luen Feng, Po-chun Hsu, Hung-yi Lee
Self-Supervised Learning (SSL) has made great strides recently. SSL speech models achieve decent performance on a wide range of downstream tasks, suggesting that they extract diffe…
Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech
Chung-Ming Chien, Jheng-Hao Lin, Chien-yu Huang +2
The few-shot multi-speaker multi-style voice cloning task is to synthesize utterances with voice and speaking style similar to a reference speaker given only a few reference sample…
Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion
Andy T. Liu, Po-chun Hsu, Hung-yi Lee
We present an unsupervised end-to-end training scheme where we discover discrete subword units from speech without using any labels. The discrete subword units are learned under an…
Rhythm-Flexible Voice Conversion without Parallel Data Using Cycle-GAN over Phoneme Posteriorgram Sequences
Cheng-chieh Yeh, Po-chun Hsu, Ju-chieh Chou +2
Speaking rate refers to the average number of phonemes within some unit time, while the rhythmic patterns refer to duration distributions for realizations of different phonemes wit…