papers

Publications (9)

cs.CL2018

Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings

Da-Rong Liu, Kuan-Yu Chen, Hung-Yi Lee +1

Unsupervised discovery of acoustic tokens from audio corpora without annotation and learning vector representations for these tokens have been widely studied. Although these techni…

cs.SD2022

Meta-TTS: Meta-Learning for Few-Shot Speaker Adaptive Text-to-Speech

Sung-Feng Huang, Chyi-Jiunn Lin, Da-Rong Liu +2

Personalizing a speech synthesis system is a highly desired application, where the system can generate speech with the user's voice with rare enrolled recordings. There are two mai…

eess.AS2022

Analyzing the Robustness of Unsupervised Speech Recognition

Guan-Ting Lin, Chan-Jan Hsu, Da-Rong Liu +2

Unsupervised speech recognition (unsupervised ASR) aims to learn the ASR system with non-parallel speech and text corpus only. Wav2vec-U has shown promising results in unsupervised…

cs.CL2021

SUPERB: Speech processing Universal PERformance Benchmark

Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang +17

Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV). The paradigm pretrains a shared model on large…

cs.SD2021

Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training

Sung-Feng Huang, Shun-Po Chuang, Da-Rong Liu +3

Speech separation has been well developed, with the very successful permutation invariant training (PIT) approach, although the frequent label assignment switching happening during…

cs.CL2020

Contextualizing ASR Lattice Rescoring with Hybrid Pointer Network Language Model

Da-Rong Liu, Chunxi Liu, Frank Zhang +3

Videos uploaded on social media are often accompanied with textual descriptions. In building automatic speech recognition (ASR) systems for videos, we can exploit the contextual in…

cs.CL2019

Completely Unsupervised Speech Recognition By A Generative Adversarial Network Harmonized With Iteratively Refined Hidden Markov Models

Kuan-Yu Chen, Che-Ping Tsai, Da-Rong Liu +2

Producing a large annotated speech corpus for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced, but collecting a rel…

cs.CL2016

Attention-based Memory Selection Recurrent Network for Language Modeling

Da-Rong Liu, Shun-Po Chuang, Hung-yi Lee

Recurrent neural networks (RNNs) have achieved great success in language modeling. However, since the RNNs have fixed size of memory, their memory cannot store all the information…

cs.CL2021

SpeechNet: A Universal Modularized Model for Speech Processing Tasks

Yi-Chen Chen, Po-Han Chi, Shu-wen Yang +7

There is a wide variety of speech processing tasks ranging from extracting content information from speech signals to generating speech signals. For different tasks, model networks…