514 citations · 543 across the 24 of their papers we have counts for
8 papers · 1 filter
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
Yingzhi Wang, Anas Alhmoud, Saad Alsahly +2
OpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-spee…
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
Fırat Öncel, Matthias Bethge, Beyza Ermis +3
In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional d…
ProGRes: Prompted Generative Rescoring on ASR n-Best
Ada Defne Tur, Adel Moumen, Mirco Ravanelli
Large Language Models (LLMs) have shown their ability to improve the performance of speech recognizers by effectively rescoring the n-best hypotheses generated during the beam sear…
Timers and Such: A Practical Benchmark for Spoken Language Understanding with Numbers
Loren Lugosch, Piyush Papreja, Mirco Ravanelli +2
This paper introduces Timers and Such, a new open source dataset of spoken English commands for common voice control use cases involving numbers. We describe the gap in existing sp…
Deep Learning for Distant Speech Recognition
Mirco Ravanelli
Deep learning is an emerging technology that is considered one of the most promising directions for reaching higher levels of artificial intelligence. Among the other achievements,…
Improving speech recognition by revising gated recurrent units
Mirco Ravanelli, Philemon Brakel, Maurizio Omologo +1
Speech recognition is largely taking advantage of deep learning, showing that substantial benefits can be obtained by modern Recurrent Neural Networks (RNNs). The most popular RNNs…