11 papers
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Albert Zeyer, Ralf Schlüter, Hermann Ney
Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal c…
LLMs and Speech: Integration vs. Combination
Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen +2
In this work, we study different approaches to utilize large language models (LLMs) for automatic speech recognition (ASR). Specifically, we compare the tight integration of an aco…
Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition
Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang +3
Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately…
Text-Utilization for Encoder-dominated Speech Recognition Models
Albert Zeyer, Tim Posielek, Ralf Schlüter +1
This paper investigates efficient methods for utilizing text-only data to improve speech recognition, focusing on encoder-dominated models that facilitate faster recognition. We pr…
Diffusion Language Models for Speech Recognition
Davyd Naveriani, Albert Zeyer, Ralf Schlüter +1
Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation.…
Reproducing and Dissecting Denoising Language Models for Speech Recognition
Dorian Koch, Albert Zeyer, Nick Rossenbach +2
Denoising language models (DLMs) have been proposed as a powerful alternative to traditional language models (LMs) for automatic speech recognition (ASR), motivated by their abilit…