3 citations · 6 across the 5 of their papers we have counts for
6 papers
Combining Contrastive and Non-Contrastive Losses for Fine-Tuning Pretrained Models in Speech Analysis
Florian Lux, Ching-Yi Chen, Ngoc Thang Vu
Embedding paralinguistic properties is a challenging task as there are only a few hours of training data available for domains such as emotional speech. One solution to this proble…
Low-Resource Multilingual and Zero-Shot Multispeaker TTS
Florian Lux, Julia Koch, Ngoc Thang Vu
While neural methods for text-to-speech (TTS) have shown great advances in modeling multiple speakers, even in zero-shot settings, the amount of data needed for those approaches is…
Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy
Sarina Meyer, Pascal Tilli, Pavel Denisov +3
In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes wit…
Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features
Florian Lux, Ngoc Thang Vu
While neural text-to-speech systems perform remarkably well in high-resource scenarios, they cannot be applied to the majority of the over 6,000 spoken languages in the world due t…
Meta-Learning for improving rare word recognition in end-to-end ASR
Florian Lux, Ngoc Thang Vu
We propose a new method of generating meaningful embeddings for speech, changes to four commonly used meta learning approaches to enable them to perform keyword spotting in continu…
ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents
Chia-Yu Li, Daniel Ortega, Dirk Väth +9
We present ADVISER - an open-source, multi-domain dialog system toolkit that enables the development of multi-modal (incorporating speech, text and vision), socially-engaged (e.g.…