5 papers · 1 filter
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Albert Zeyer, Ralf Schlüter, Hermann Ney
Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal c…
Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition
Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang +3
Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately…
Text-Utilization for Encoder-dominated Speech Recognition Models
Albert Zeyer, Tim Posielek, Ralf Schlüter +1
This paper investigates efficient methods for utilizing text-only data to improve speech recognition, focusing on encoder-dominated models that facilitate faster recognition. We pr…
Diffusion Language Models for Speech Recognition
Davyd Naveriani, Albert Zeyer, Ralf Schlüter +1
Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation.…
Dynamic Acoustic Model Architecture Optimization in Training for ASR
Jingjing Xu, Zijian Yang, Albert Zeyer +3
Architecture design is inherently complex. Existing approaches rely on either handcrafted rules, which demand extensive empirical expertise, or automated methods like neural archit…