6 papers
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
Pehuén Moure, Pehuén Moure, Niclas Pokel +5
Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility of improving performance by co…
Demonstration of Adapt4Me: An Uncertainty-Aware Authoring Environment for Personalizing Automatic Speech Recognition to Non-normative Speech
Niclas Pokel, Yiming Zhao, Pehuén Moure +2
Personalizing Automatic Speech Recognition (ASR) for non-normative speech remains challenging because data collection is labor-intensive and model training is technically complex.…
Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition
Niclas Pokel, Pehuén Moure, Roman Boehringer +2
Speech impairments resulting from congenital disorders, such as cerebral palsy, down syndrome, or apert syndrome, as well as acquired brain injuries due to stroke, traumatic accide…
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
Niclas Pokel, Pehuén Moure, Roman Böhringer +1
ASR systems struggle with non-normative speech due to high acoustic variability and data scarcity. We propose a data-efficient method using phoneme-level uncertainty to guide fine-…
Reasoning aligns language models to human cognition
Gonçalo Guiomar, Elia Torre, Pehuen Moure +4
Do language models make decisions under uncertainty like humans do, and what role does chain-of-thought (CoT) reasoning play in the underlying decision process? We introduce an act…
Adapting Foundation Speech Recognition Models to Impaired Speech: A Semantic Re-chaining Approach for Personalization of German Speech
Niclas Pokel, Pehuén Moure, Roman Boehringer +1
Speech impairments caused by conditions such as cerebral palsy or genetic disorders pose significant challenges for automatic speech recognition (ASR) systems. Despite recent advan…