activity
20152021
most citedImproved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering

3 citations · 11 across the 12 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD20193 cited

Interrupted and cascaded permutation invariant training for speech separation

Gene-Ping Yang, Szu-Lin Wu, Yao-Wen Mao +2

Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the mini…

cs.SD20193 cited

Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering

Gene-Ping Yang, Chao-I Tuan, Hung-Yi Lee +1

Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the stan…

cs.SD2018

Rhythm-Flexible Voice Conversion without Parallel Data Using Cycle-GAN over Phoneme Posteriorgram Sequences

Cheng-chieh Yeh, Po-chun Hsu, Ju-chieh Chou +2

Speaking rate refers to the average number of phonemes within some unit time, while the rhythmic patterns refer to duration distributions for realizations of different phonemes wit…

cs.SD2018

Transcribing Lyrics From Commercial Song Audio: The First Step Towards Singing Content Processing

Che-Ping Tsai, Yi-Lin Tuan, Lin-shan Lee

Spoken content processing (such as retrieval and browsing) is maturing, but the singing content is still almost completely left out. Songs are human voice carrying plenty of semant…

cs.SD20171 cited

Personalized Acoustic Modeling by Weakly Supervised Multi-Task Deep Learning using Acoustic Tokens Discovered from Unlabeled Data

Cheng-Kuan Wei, Cheng-Tao Chung, Hung-Yi Lee +1

It is well known that recognizers personalized to each user are much more effective than user-independent recognizers. With the popularity of smartphones today, although it is not…

cs.SD2016

Audio Word2Vec: Unsupervised Learning of Audio Segment Representations using Sequence-to-sequence Autoencoder

Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen +2

The vector representations of fixed dimensionality for words (in text) offered by Word2Vec have been shown to be very useful in many application scenarios, in particular due to the…