195 citations · 197 across the 2 of their papers we have counts for
6 papers
REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling
Hu Hu, Xuesong Yang, Zeynab Raeesy +6
Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…
AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
Kaizhi Qian, Yang Zhang, Shiyu Chang +2
Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas. Deep style transfer algorithms, such as generative adversarial netwo…
When CTC Training Meets Acoustic Landmarks
Di He, Xuesong Yang, Boon Pang Lim +3
Connectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but s…
Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks
Di He, Boon Pang Lim, Xuesong Yang +2
Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cue…
Deep Learning Based Speech Beamforming
Kaizhi Qian, Yang Zhang, Shiyu Chang +3
Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the sp…
Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg +3
The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significa…