activity
20242026
collaborators

8 papers

cs.SD2026

Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0

Natalie Engert, Dominik Wagner, Korbinian Riedhammer +1

Wav2vec 2.0 (W2V2) has shown strong performance in pathological speech analysis by effectively capturing the characteristics of atypical speech. Despite its success, it remains unc…

eess.AS2025

On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts

Kashaf Gulzar, Dominik Wagner, Sebastian P. Bayerl +3

Automatic transcription of stuttered speech remains a challenge, even for modern end-to-end (E2E) automatic speech recognition (ASR) frameworks. Dysfluencies and fluency-shaping ar…

cs.SD2025

Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks

Dominik Wagner, Ilja Baumann, Tobias Bocklet

Cycle-consistent generative adversarial networks have been widely used in non-parallel voice conversion (VC). Their ability to learn mappings between source and target features wit…

cs.SD2025

Optimized Self-supervised Training with BEST-RQ for Speech Recognition

Ilja Baumann, Dominik Wagner, Korbinian Riedhammer +1

Self-supervised learning has been successfully used for various speech related tasks, including automatic speech recognition. BERT-based Speech pre-Training with Random-projection…

cs.LG2024

Optimized Speculative Sampling for GPU Hardware Accelerators

Dominik Wagner, Seanie Lee, Ilja Baumann +3

In this work, we optimize speculative sampling for parallel hardware accelerators to improve sampling speed. We notice that substantial portions of the intermediate matrices necess…

cs.CL2024

MMUTF: Multimodal Multimedia Event Argument Extraction with Unified Template Filling

Philipp Seeberger, Dominik Wagner, Korbinian Riedhammer

With the advancement of multimedia technologies, news documents and user-generated content are often represented as multiple modalities, making Multimedia Event Extraction (MEE) an…