activity
20162022
most citedSyntactic and Semantic Features For Code-Switching Factored Language Models

70 citations · 105 across the 23 of their papers we have counts for

collaborators

44 papers

cs.CL20222 cited

ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic - English

Injy Hamed, Nizar Habash, Slim Abdennadher +1

We present our work on collecting ArzEn-ST, a code-switched Egyptian Arabic - English Speech Translation Corpus. This corpus is an extension of the ArzEn speech corpus, which was c…

cs.CL2022

Combining Contrastive and Non-Contrastive Losses for Fine-Tuning Pretrained Models in Speech Analysis

Florian Lux, Ching-Yi Chen, Ngoc Thang Vu

Embedding paralinguistic properties is a challenging task as there are only a few hours of training data available for domains such as emotional speech. One solution to this proble…

cs.CL20222 cited

Low-Resource Multilingual and Zero-Shot Multispeaker TTS

Florian Lux, Julia Koch, Ngoc Thang Vu

While neural methods for text-to-speech (TTS) have shown great advances in modeling multiple speakers, even in zero-shot settings, the amount of data needed for those approaches is…

cs.CL2022

Improving Semi-supervised End-to-end Automatic Speech Recognition using CycleGAN and Inter-domain Losses

Chia-Yu Li, Ngoc Thang Vu

We propose a novel method that combines CycleGAN and inter-domain losses for semi-supervised end-to-end automatic speech recognition. Inter-domain loss targets the extraction of an…

cs.SD20223 cited

Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy

Sarina Meyer, Pascal Tilli, Pavel Denisov +3

In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes wit…

cs.CL2022

BPE vs. Morphological Segmentation: A Case Study on Machine Translation of Four Polysynthetic Languages

Manuel Mager, Arturo Oncevay, Elisabeth Mager +2

Morphologically-rich polysynthetic languages present a challenge for NLP systems due to data sparsity, and a common strategy to handle this issue is to apply subword segmentation.…