most citedGenerative Models for Improved Naturalness, Intelligibility, and Voicing of Whispered Speech

6 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2024

Large Language Models for Dysfluency Detection in Stuttered Speech

Dominik Wagner, Sebastian P. Bayerl, Ilja Baumann +3

Accurately detecting dysfluencies in spoken language can help to improve the performance of automatic speech and language processing components and support the development of more…

cs.SD2024

Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models

Dominik Wagner, Ilja Baumann, Korbinian Riedhammer +1

This paper explores the improvement of post-training quantization (PTQ) after knowledge distillation in the Whisper speech foundation model family. We address the challenge of outl…

cs.CV2024

Semmeldetector: Application of Machine Learning in Commercial Bakeries

Thomas H. Schmitt, Maximilian Bundscherer, Tobias Bocklet

The Semmeldetector, is a machine learning application that utilizes object detection models to detect, classify and count baked goods in images. Our application allows commercial b…

eess.AS2023

A Stutter Seldom Comes Alone -- Cross-Corpus Stuttering Detection as a Multi-label Problem

Sebastian P. Bayerl, Dominik Wagner, Ilja Baumann +4

Most stuttering detection and classification research has viewed stuttering as a multi-class classification problem or a binary detection task for each dysfluency type; however, th…

cs.SD20236 cited

Generative Models for Improved Naturalness, Intelligibility, and Voicing of Whispered Speech

Dominik Wagner, Sebastian P. Bayerl, Hector A. Cordourier Maruri +1

This work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal…