9 citations · 19 across the 8 of their papers we have counts for
6 papers · 1 filter
Speech-text based multi-modal training with bidirectional attention for improved speech recognition
Yuhang Yang, Haihua Xu, Hao Huang +2
To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the s…
Minimum word error training for non-autoregressive Transformer-based code-switching ASR
Yizhou Peng, Jicheng Zhang, Haihua Xu +2
Non-autoregressive end-to-end ASR framework might be potentially appropriate for code-switching recognition task thanks to its inherent property that present output token being ind…
E2E-based Multi-task Learning Approach to Joint Speech and Accent Recognition
Jicheng Zhang, Yizhou Peng, Pham Van Tung +3
In this paper, we propose a single multi-task learning framework to perform End-to-End (E2E) speech recognition (ASR) and accent recognition (AR) simultaneously. The proposed frame…
The NTU-AISG Text-to-speech System for Blizzard Challenge 2020
Haobo Zhang, Tingzhi Mao, Haihua Xu +1
We report our NTU-AISG Text-to-speech (TTS) entry systems for the Blizzard Challenge 2020 in this paper. There are two TTS tasks in this year's challenge, one is a Mandarin TTS tas…
Monolingual Data Selection Analysis for English-Mandarin Hybrid Code-switching Speech Recognition
Haobo Zhang, Haihua Xu, Van Tung Pham +2
In this paper, we conduct data selection analysis in building an English-Mandarin code-switching (CS) speech recognition (CSSR) system, which is aimed for a real CSSR contest in Ch…
Approaches to Improving Recognition of Underrepresented Named Entities in Hybrid ASR Systems
Tingzhi Mao, Yerbolat Khassanov, Van Tung Pham +3
In this paper, we present a series of complementary approaches to improve the recognition of underrepresented named entities (NE) in hybrid ASR systems without compromising overall…