works on

From the 2 of 20 linked papers with an AI index.

collaborators

20 papers

eess.AS2026

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

Haolin He, Renhe Sun, Zheqi Dai +16

DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Au…

eess.AS2026

Efficient Text-to-Audio Generation via Pruning

Arshdeep Singh, Yi Yuan, Yun Chen +2

The paper applies filter‑based pruning to the U‑Net backbone of the AudioLDM text‑to‑audio diffusion model, reducing most of its parameters and compute while preserving generation…

cs.SD2026

Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation

Paul A. Bereuter, Mark D. Plumbley, Alois Sontacchi

The paper adapts pretrained speech enhancement models to perform singing voice separation using fine‑tuning and low‑rank adaptation, achieving better separation with limited singin…

eess.AS2026

Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

Yanze Xu, Wenwu Wang, Mark D. Plumbley

Neural networks can be trained to learn task-relevant representations from data. Understanding how these networks make decisions falls within the Explainable AI (XAI) domain. This…

eess.AS2026

Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation

Yanze Xu, Mark D. Plumbley, Wenwu Wang

Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of expla…

cs.SD2026

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

Yi Yuan, Xubo Liu, Haohe Liu +5

With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite produc…