6 papers · 1 filter
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
Dominik Wagner, Ilja Baumann, Tobias Bocklet
Cycle-consistent generative adversarial networks have been widely used in non-parallel voice conversion (VC). Their ability to learn mappings between source and target features wit…
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
Dominik Wagner, Ilja Baumann, Natalie Engert +4
In this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent au…
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
Ilja Baumann, Dominik Wagner, Korbinian Riedhammer +1
Self-supervised learning has been successfully used for various speech related tasks, including automatic speech recognition. BERT-based Speech pre-Training with Random-projection…
Large Language Models for Dysfluency Detection in Stuttered Speech
Dominik Wagner, Sebastian P. Bayerl, Ilja Baumann +3
Accurately detecting dysfluencies in spoken language can help to improve the performance of automatic speech and language processing components and support the development of more…
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
Dominik Wagner, Ilja Baumann, Korbinian Riedhammer +1
This paper explores the improvement of post-training quantization (PTQ) after knowledge distillation in the Whisper speech foundation model family. We address the challenge of outl…
A Survey of Music Generation in the Context of Interaction
Ismael Agchar, Ilja Baumann, Franziska Braun +4
In recent years, machine learning, and in particular generative adversarial neural networks (GANs) and attention-based neural networks (transformers), have been successfully used t…