activity
20242026
collaborators

5 papers

cs.CL2026

ALAS: An Automatic Latent Alignment Score for Audio Language Models

Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli +1

Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) beha…

cs.CL2025

Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down

Yingzhi Wang, Anas Alhmoud, Saad Alsahly +2

OpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-spee…

cs.SD2025

What Are They Doing? Joint Audio-Speech Co-Reasoning

Yingzhi Wang, Pooneh Mousavi, Artem Ploujnikov +1

In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Audito…

cs.CL2024

Open Universal Arabic ASR Leaderboard

Yingzhi Wang, Anas Alhmoud, Muhammad Alqurishi

In recent years, the enhanced capabilities of ASR models and the emergence of multi-dialect datasets have increasingly pushed Arabic ASR model development toward an all-dialect-in-…

cs.LG2024

Open-Source Conversational AI with SpeechBrain 1.0

Mirco Ravanelli, Titouan Parcollet, Adel Moumen +30

SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker re…