From the 1 of 21 linked papers with an AI index.
21 papers
FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation
Yi Yuan, Xubo Liu, Haohe Liu +3
Language-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable inte…
Efficient Text-to-Audio Generation via Pruning
Arshdeep Singh, Yi Yuan, Yun Chen +2
The paper applies filter‑based pruning to the U‑Net backbone of the AudioLDM text‑to‑audio diffusion model, reducing most of its parameters and compute while preserving generation…
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
Yanze Xu, Wenwu Wang, Mark D. Plumbley
Neural networks can be trained to learn task-relevant representations from data. Understanding how these networks make decisions falls within the Explainable AI (XAI) domain. This…
Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation
Yanze Xu, Mark D. Plumbley, Wenwu Wang
Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of expla…
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
Yi Yuan, Xubo Liu, Haohe Liu +5
With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite produc…
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
Junqi Zhao, Chenxing Li, Jinzheng Zhao +4
We propose a general feedback-driven retrieval-augmented generation (RAG) approach that leverages Large Audio Language Models (LALMs) to address the missing or imperfect synthesis…