activity
20202022
most citedAnonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy

3 citations · 6 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2022

Combining Contrastive and Non-Contrastive Losses for Fine-Tuning Pretrained Models in Speech Analysis

Florian Lux, Ching-Yi Chen, Ngoc Thang Vu

Embedding paralinguistic properties is a challenging task as there are only a few hours of training data available for domains such as emotional speech. One solution to this proble…

cs.CL20222 cited

Low-Resource Multilingual and Zero-Shot Multispeaker TTS

Florian Lux, Julia Koch, Ngoc Thang Vu

While neural methods for text-to-speech (TTS) have shown great advances in modeling multiple speakers, even in zero-shot settings, the amount of data needed for those approaches is…

cs.SD20223 cited

Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy

Sarina Meyer, Pascal Tilli, Pavel Denisov +3

In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes wit…

cs.CL2022

Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features

Florian Lux, Ngoc Thang Vu

While neural text-to-speech systems perform remarkably well in high-resource scenarios, they cannot be applied to the majority of the over 6,000 spoken languages in the world due t…

eess.AS2021

Meta-Learning for improving rare word recognition in end-to-end ASR

Florian Lux, Ngoc Thang Vu

We propose a new method of generating meaningful embeddings for speech, changes to four commonly used meta learning approaches to enable them to perform keyword spotting in continu…

cs.CL20201 cited

ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents

Chia-Yu Li, Daniel Ortega, Dirk Väth +9

We present ADVISER - an open-source, multi-domain dialog system toolkit that enables the development of multi-modal (incorporating speech, text and vision), socially-engaged (e.g.…