70 citations · 108 across the 28 of their papers we have counts for
41 papers · 1 filter
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
Chia-Yu Li, Ngoc Thang Vu
Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial a…
Neural Machine Translation for the Indigenous Languages of the Americas: An Introduction
Manuel Mager, Rajat Bhatnagar, Graham Neubig +2
Neural models have drastically advanced state of the art for machine translation (MT) between high-resource languages. Traditionally, these models rely on large amounts of training…
Oh, Jeez! or Uh-huh? A Listener-aware Backchannel Predictor on ASR Transcriptions
Daniel Ortega, Chia-Yu Li, Ngoc Thang Vu
This paper presents our latest investigation on modeling backchannel in conversations. Motivated by a proactive backchanneling theory, we aim at developing a system which acts as a…
Modeling Speaker-Listener Interaction for Backchannel Prediction
Daniel Ortega, Sarina Meyer, Antje Schweitzer +1
We present our latest findings on backchannel modeling novelly motivated by the canonical use of the minimal responses Yeah and Uh-huh in English and their correspondent tokens in…
ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic - English
Injy Hamed, Nizar Habash, Slim Abdennadher +1
We present our work on collecting ArzEn-ST, a code-switched Egyptian Arabic - English Speech Translation Corpus. This corpus is an extension of the ArzEn speech corpus, which was c…
Combining Contrastive and Non-Contrastive Losses for Fine-Tuning Pretrained Models in Speech Analysis
Florian Lux, Ching-Yi Chen, Ngoc Thang Vu
Embedding paralinguistic properties is a challenging task as there are only a few hours of training data available for domains such as emotional speech. One solution to this proble…