1 citations · 1 across the 4 of their papers we have counts for
9 papers
Investigating Faithfulness in Large Audio Language Models
Pooneh Mousavi, Lovenya Jain, Mirco Ravanelli +1
Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks. While these models can generate Ch…
ALAS: An Automatic Latent Alignment Score for Audio Language Models
Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli +1
Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) beha…
DASB - Discrete Audio and Speech Benchmark
Pooneh Mousavi, Jarod Duret, Darius Petermann +5
Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both gen…
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
Jihoon Jeong, Pooneh Mousavi, Mirco Ravanelli +1
Large audio-language models (LALMs) can generate reasoning chains for their predictions, but it remains unclear whether these reasoning chains remain grounded in the input audio. I…
Discrete Audio Tokens: More Than a Survey!
Pooneh Mousavi, Gallil Maimon, Adel Moumen +18
Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and infere…
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
Pooneh Mousavi, Shubham Gupta, Cem Subakan +1
Foundation models based on large language models (LLMs) have shown great success in handling various tasks and modalities. However, adapting these models for general-purpose audio-…