4 papers · 1 filter
LLMs and Speech: Integration vs. Combination
Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen +2
In this work, we study different approaches to utilize large language models (LLMs) for automatic speech recognition (ASR). Specifically, we compare the tight integration of an aco…
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
Peter Vieting, Maximilian Kannen, Benedikt Hilmes +2
Neural front-ends are an appealing alternative to traditional, fixed feature extraction pipelines for automatic speech recognition (ASR) systems since they can be directly trained…
Unified Learnable 2D Convolutional Feature Extraction for ASR
Peter Vieting, Benedikt Hilmes, Ralf Schlüter +1
Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for dif…
The Conformer Encoder May Reverse the Time Dimension
Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen +2
We sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, Further investigation shows that t…