activity
20172019
most citedOrder-Free RNN with Visual Attention for Multi-Label Classification

52 citations · 53 across the 2 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20191 cited

From Semi-supervised to Almost-unsupervised Speech Recognition with Very-low Resource by Jointly Learning Phonetic Structures from Audio and Text Embeddings

Yi-Chen Chen, Sung-Feng Huang, Hung-yi Lee +1

Producing a large amount of annotated speech data for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced. However, we…

cs.CL2018

Improved Audio Embeddings by Adjacency-Based Clustering with Applications in Spoken Term Detection

Sung-Feng Huang, Yi-Chen Chen, Hung-yi Lee +1

Embedding audio signal segments into vectors with fixed dimensionality is attractive because all following processing will be easier and more efficient, for example modeling, class…

cs.CL2018

Almost-unsupervised Speech Recognition with Close-to-zero Resource Based on Phonetic Structures Learned from Very Small Unpaired Speech and Text Data

Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang +2

Producing a large amount of annotated speech data for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced. However, we…

cs.CL2018

Phonetic-and-Semantic Embedding of Spoken Words with Applications in Spoken Content Retrieval

Yi-Chen Chen, Sung-Feng Huang, Chia-Hao Shen +2

Word embedding or Word2Vec has been successful in offering semantics for text words learned from the context of words. Audio Word2Vec was shown to offer phonetic structures for spo…

cs.CL2018

Towards Unsupervised Automatic Speech Recognition Trained by Unaligned Speech and Text only

Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang +1

Automatic speech recognition (ASR) has been widely researched with supervised approaches, while many low-resourced languages lack audio-text aligned data, and supervised methods ca…