4 papers
ALAS: An Automatic Latent Alignment Score for Audio Language Models
Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli +1
Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) beha…
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
Yingzhi Wang, Anas Alhmoud, Saad Alsahly +2
OpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-spee…
What Are They Doing? Joint Audio-Speech Co-Reasoning
Yingzhi Wang, Pooneh Mousavi, Artem Ploujnikov +1
In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Audito…
Open Universal Arabic ASR Leaderboard
Yingzhi Wang, Anas Alhmoud, Muhammad Alqurishi
In recent years, the enhanced capabilities of ASR models and the emergence of multi-dialect datasets have increasingly pushed Arabic ASR model development toward an all-dialect-in-…