13 papers
Phoneme-First Prediction for LLM-Based Speech Recognition
Jakob Poncelet, Hugo Van hamme
Recent research has explored integrating Large Language Models (LLMs) with speech encoders to create speech-augmented LLMs capable of contextualized speech recognition. The main ch…
Speech Encoder Fusion for LLM-based Automatic Speech Recognition
Jakob Poncelet, Hugo Van hamme
Speech-aware large language models (LLMs) can incorporate speech through pre-trained acoustic encoders that project speech features into the LLM embedding space. While the choice o…
Towards Deep Contextual Reasoning from Broad Descriptions for ASR with Speech-LLM via Metadata-Driven Reasoning Chains
Jakob Poncelet, Hugo Van hamme
Speech recognition often fails on rare, domain-specific terms and context-related named entities. Existing contextualization techniques typically bias decoding with keywords or phr…
Parameter-Efficient Continual Learning for Automatic Speech Recognition
Steven Vander Eeckt, Hugo Van hamme
Speech foundation models enable strong general-purpose ASR and are attractive for downstream adaptation. However, their size and the catastrophic forgetting induced by sequential f…
GLoRIA: Gated Low-Rank Interpretable Adaptation for Dialectal ASR
Pouya Mehralian, Melissa Farasyn, Anne Breitbarth +2
Automatic Speech Recognition (ASR) in dialect-heavy settings remains challenging due to strong regional variation and limited labeled data. We propose GLoRIA, a parameter-efficient…
Inverse-Hessian Regularization for Continual Learning in ASR
Steven Vander Eeckt, Hugo Van hamme
Catastrophic forgetting remains a major challenge for continual learning (CL) in automatic speech recognition (ASR), where models must adapt to new domains without losing performan…