4 papers
Phoneme-First Prediction for LLM-Based Speech Recognition
Jakob Poncelet, Hugo Van hamme
Recent research has explored integrating Large Language Models (LLMs) with speech encoders to create speech-augmented LLMs capable of contextualized speech recognition. The main ch…
Speech Encoder Fusion for LLM-based Automatic Speech Recognition
Jakob Poncelet, Hugo Van hamme
Speech-aware large language models (LLMs) can incorporate speech through pre-trained acoustic encoders that project speech features into the LLM embedding space. While the choice o…
Towards Deep Contextual Reasoning from Broad Descriptions for ASR with Speech-LLM via Metadata-Driven Reasoning Chains
Jakob Poncelet, Hugo Van hamme
Speech recognition often fails on rare, domain-specific terms and context-related named entities. Existing contextualization techniques typically bias decoding with keywords or phr…
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
Jakob Poncelet, Hugo Van hamme
The recent advancement of speech recognition technology has been driven by large-scale datasets and attention-based architectures, but many challenges still remain, especially for…